En
← All articles

SenseTime Announces SenseNova U1 Pro: Multimodal Model Competition Shifts to Delivering Complex Tasks

SenseTime unveiled SenseNova U1 Pro at the 2026 World Artificial Intelligence Conference, positioning it as a native multimodal agent foundation for…

AuthorOpen Market NotesTypeArticle
SenseTime Announces SenseNova U1 Pro: Multimodal Model Competition Shifts to Delivering Complex Tasks

1. What happened

On July 18, at the 2026 World Artificial Intelligence Conference (WAIC), SenseTime Technology officially launched SenseNova U1 Pro, the flagship product in its large model family. SenseTime positions it as a “delivery-grade” native multimodal agent foundation for long-horizon tasks, aiming for the model to go beyond one-shot content generation and participate in understanding, planning, execution, inspection, and revision for complex tasks.

SenseNova U1 Pro is the flagship version of the “SenseNova U” series, covering scenarios such as visual design, infographics, presentation materials, video storyboards, and scientific illustrations. According to SenseTime, the preview version has already begun invite-only testing, and the official version is scheduled for public release in August 2026, along with pricing plans and API services.

2. Why it matters

From content generation to deliverables

Traditional generative models have usually focused on producing images or other content from a single prompt, but with U1 Pro, SenseTime wants to handle longer-process tasks. The model needs to understand complex objectives, continue multiple rounds of creation and revision, and ultimately deliver a visually polished output.

SenseTime summarizes this direction as a shift from “Generation” to “Delivery,” drawing an analogy to the evolution in coding from Copilot and Vibe Coding to Agentic Coding. If this product direction is implemented reliably, competition among visual models will extend beyond the quality of a single image to include information organization, style consistency, detail refinement, and task completion capability.

Visual models are beginning to enter production workflows

U1 Pro is aimed not at entertainment image generation alone, but at workflows such as enterprise infographics and PPTs, commercial posters, e-commerce product visuals, scientific illustrations, video concept design, and anime storyboarding. This product positioning shows that SenseTime is trying to embed multimodal models further into office work, commercial design, and content production processes.

That said, the information currently available is mainly based on SenseTime’s presentation materials and media reports, and it is still insufficient to independently verify stability and efficiency across different real-world workflows.

3. Public information and product capabilities

According to SenseTime, SenseNova U1 Pro mainly offers the following capabilities.

  • Professional visual design: Emphasizes composition, color, and layout effects, aiming to reduce the “AI feel” often seen in traditional generated images.
  • Native 8K output: Supports native resolutions of up to 8K, intended for ultra-long, ultra-large, and high-information-density visual content production.
  • Image-text and detail control: Enhances handling of relationships among text, images, layout, and details, making it suitable for infographics, presentation materials, and complex knowledge diagrams.
  • Long-term Agentic closed loop: For complex objectives, it can run dozens of agent generation cycles and edit both the overall style and local text at the same time.

Examples shared by SenseTime include first designing the world-building, character settings, costumes and weapons, and environmental color palette for a video story, and then generating up to 22 logically consistent consecutive storyboards. Another example integrates promotion paths in a sports match, player matchup relationships, tactical systems, passing networks, and ball-touch heat maps, then outputs them as a high-definition chart report.

SenseTime also presented a 4:1 ultra-wide image work featuring urban landmarks and dense text information, demonstrating the model’s capabilities in long canvases, complex layouts, and image-text integration.

4. Technical route and evidence limits

SenseTime says U1 Pro is built on the NEO-unify core architecture, unifying language representation and visual representation within a native multimodal system in an effort to reduce the gap between understanding and generation.

SenseTime also says its open-source SenseNova-Vision unified visual large model brings traditional vision tasks such as instance segmentation and object detection into the capabilities of a general-purpose large model, and that this will strengthen multimodal agents’ understanding of the visual world.

At present, the publicly available materials do not disclose U1 Pro’s parameter scale, training data, inference cost, context length, or a full technical report. Capabilities such as native 8K, a low text error rate, and dozens of agent loops are mainly based on SenseTime’s announcements and media paraphrases, and still require further verification through public samples, third-party benchmarks, and real user feedback.

5. What to watch next

  1. Whether the official version is released as planned: The preview version has already begun invite-only testing, and the official version is scheduled for release in August 2026. The actual launch timing and scope still need to be confirmed.
  2. API and pricing: SenseTime plans to provide the corresponding API services, and billing method, call limits, inference speed, and 8K output costs will directly affect enterprise adoption.
  3. Third-party evaluation results: Text accuracy, complex layouts, long-task completion, style consistency, and independent performance under multiple rounds of revision still need verification.
  4. Differences from other products in the series: The differences in capability, speed, price, and applicable scenarios between U1 Pro and SenseNova U1 and U1-Fast may influence user choice.
  5. Commercialization progress: SenseTime says the relevant capabilities have been validated in products such as “Xiaohuanxiong” and “Seko,” but it has not disclosed specific call volumes, customer retention, or revenue contribution at this stage.

Conclusion

The launch of SenseNova U1 Pro shows that competition in multimodal models is expanding from the generation quality of a single image to the planning of complex tasks, iteration, and delivery of deliverables. Through native 8K, image-text control, and long-term agent loops, SenseTime is trying to push visual models from creative tools toward a more complete productivity system.

Whether this positioning holds will ultimately depend on the official release, API pricing, third-party evaluations, and actual user feedback.

Original source

36氪

正文图片
正文图片

Information only. No investment, legal, tax, or financial advice.