Back to feed

From One Prompt to a Working Vision Agent: What NVIDIA Cosmos and VSS 3.3 Change

The livestream on building a vision analytics agent from a single natural-language prompt demos search, alerting and bottle measurement in a juice factory with NVIDIA Cosmos and VSS 3.3.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — PQJKs1dyK7I
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

Building a full video analytics agent that watches cameras, scans archives and raises alerts from a single-sentence prompt sounds like a bold claim, and the October 1 livestream sets out to prove exactly that on stage. The host welcomes Adam from VSS product management and Hassan from the Metropolis technical team, framing the broadcast as the first hands-on preview of the VSS 3.3 novelties ahead of GTC Berlin. One idea carries the whole session: the developer describes what they want in natural language, and a coding agent turns it into a running application. The announcement and viewing link for this broadcast were posted in the official thread on forums.developer.nvidia.com, where the event is explicitly billed as a herald of the GTC Berlin program.

VSS, short for Video Search and Summarization, is presented as a reference architecture under the NVIDIA Metropolis umbrella that takes video analytics end to end. The system ingests massive piles of live or archived video and turns them into natural-language search, visual question answering, verified alerts and automated reporting. Under the hood, vision-language models, large language models such as Nemotron, retrieval-augmented generation and tools connected through the Model Context Protocol work together. Because a typical production agent bundles ingestion, detection, alerting, search, summarization and reporting under one roof, this integrity matters a great deal. The current capability list and the cloud trial link live on the official Blueprint card at build.nvidia.com, which is recommended as the first stop for anyone wanting to try the architecture.

At the heart of the visual intelligence sits Cosmos, positioned in the broadcast as an open vision-language reasoning model. The critical distinction is a two-stage split: first an embedding model finds candidate clips in the archive, then Cosmos verifies whether the requested visual condition truly occurred. Since looking relevant and actually containing the requested event are different things, this retrieve-then-verify split cuts false alarms substantially. On the Cosmos 3 side, a streaming NIM capability is announced that replaces chunked processing with a rolling window evaluating every frame in turn. The Cosmos 3 family design, unifying reasoning, world generation and action generation in one model, is explained in detail in the technical blog post on developer.nvidia.com, where open models and training recipes are shared.

The architecture divides into three layers, each with a distinct responsibility. The real-time layer performs feature extraction, embedding generation and stream understanding, publishing results to a message broker. The analytics layer enriches this data into trajectories, incidents and verified alert records. The agent layer orchestrates search, question answering, summarization and clip retrieval tools through the Model Context Protocol. Shared infrastructure includes the video input-output and storage service, Kafka, Redis, Elasticsearch and an ingress load balancer, and the skill system reuses them instead of installing them from scratch each time. The service references and configuration options for this layered design are published in the official VSS documentation at docs.nvidia.com, which also explains the deployment profiles.

The recorder path and the live path are deliberately separated, and this split drives the entire demo. The recorder path handles storage, indexing and retrieval while the live path continuously watches the stream against the operator-defined condition. Rule management is collected in a component called the alert bridge, which guarantees the rule is received correctly while Cosmos evaluates the stream. When an event is caught, the application exposes it together with its video evidence so users can see what happened with their own eyes. Because this design separates continuous monitoring from rule management, the system stays both understandable and easy to maintain.

The operator window onto the world has two parts: the VSS interface and chat connected through OpenClaw. In the demo, an assistant named Astra does the conversational work, running inside a coding agent called Codex with the Build Vision Agent skill. The tool heard as Codeex in the video is actually the Codex coding agent, and the local model heard as Neatron is actually the Nemotron model. Model responsibilities are divided cleanly: the embedding model finds the relevant video, the vision-language model evaluates the visuals, and the language model carries the conversation with the operator. The operator question itself is strikingly plain: show liquid leaking from the bottles.

After the abstract architecture, the livestream steers toward a concrete business: monitoring an orange juice factory. Two basic operator needs are defined up front: finding what happened in recorded footage and getting notified the moment it happens again on a live camera. These map directly onto the where-did-it-happen versus tell-me-when-it-recurs distinction, which justifies the dual-path architecture. The original factory footage serves the search and alerting example while bottling line footage is reserved for the later measurement demo. For the live side, recorded footage is replayed over RTSP so streaming workflows can be exercised with repeatable video.

In the first prompt the developer describes the use case, and the skill responds by asking about capabilities and inputs before proposing an architecture. From the ready-made profiles, the custom configuration combining search with real-time alerting is selected, and the proposal lands on screen: real-time alerts on top of the Foundation, keeping the web UI and operator chat. The workload is split across GPUs: Cosmos 3 runs on one card, the language model on another, and the RT-DETR detector on separate resources. The alert condition is defined in a sentence: visible juice escaping from bottles during filling. Deployment never starts before the developer approves the architecture, and with models pre-downloaded the session time goes to the application itself.

On the recorded-video side the data flow runs in traceable steps. The video input-output and storage service provides access to footage, the embedding model turns visual content into searchable representations, and Elasticsearch holds the search index. When a question arrives, the system first gathers candidate clips, then Cosmos evaluates each candidate against the requested visual condition one by one. The gap between looking relevant and truly containing the event closes at this point. Because this pipeline meets alerts and chat in the same interface, the operator scans history and watches the live feed from one screen. Uniting search and monitoring in one application removes the cost of building two separate systems.

Live demo at a juice factory

Once search and alerting are combined, the demo stretches toward a third capability: per-bottle measurement. For a factory operator, how high the liquid reached in each bottle and how that compares with the expected level matters at least as much as alarms. For this extension, RF-DETR based segmentation masks come into play: one marks the bottle and another marks the visible liquid, pixel by pixel. Because true pixel regions replace bounding boxes, measurement becomes far more precise. The masks come from the GPU-running model, and the measurement logic combines them with camera calibration to estimate the visible liquid height inside the bottle.

Two different time scales of measurement are stressed, and this distinction shapes the operator experience. While filling continues, the displayed number counts as provisional and a half-filled bottle never gets stamped as an underfill. Once the cycle completes, the component stores the final result with the bottle identity, enabling retrospective comparison. Cycles 8126 and 8127 are shown as examples of such identified records. When the view is occluded or the measurement is weak, the interface shows uncertainty openly instead of faking confidence. This honesty matters because the operator must know how far to trust each number.

When records complete, every bottle appears with its own identity, measurement and verdict, and bottle 8105 steals the show as a 28 percent underfill. The operator can attach this record straight into the VSS chat and ask what happened to it, and the agent produces a concrete explanation from the measurements returned by the filling tools. The explanation can be checked against the interface value side by side, and follow-up questions can continue on the same record. Two bottles can be attached together and their differences interrogated. The operator thus reaches conclusions by talking to the agent instead of memorizing every field, escaping the burden of remembering what happened five bottles ago on the live preview.

The system does not only look at cameras; additional sensor streams can be connected when needed, and the agent folds these readings into its decisions. Calibration is flagged as mandatory in production: the demo used a sample calibration teaching where 100 percent sits on these bottles, but a real line must apply all existing calibration procedures in full. The camera-relative percent unit is explained with care: the on-screen figure is not millimeters or physical volume but the ratio of visible liquid height to the bottle. This definitional clarity prevents misreading numbers across different lines and camera setups. Reference values and lower limits are configured against the same ratio. Digital twin integration lightens this load by enabling advance calibration in a virtual environment: twins built with Omniverse let teams test settings before touching the physical line, cutting commissioning risk and production loss, and as the virtual-real gap closes, calibration reliability improves while multi-camera plants save serious time with virtual preparation instead of on-site tuning.

Cost math: 3 dollars to build, pennies to run

The cost section runs on a single RTX Pro GPU with cloud pricing, answering two questions: what does it cost to build, and what does it cost to keep running. The first agent stands up after about 30 minutes of prompting at a stated cost of 3 dollars. Further time for fine-tuning and extension is recommended after setup, but the first working agent connected to the system comes at that budget. A tidy summary of this analysis and the technical background of both new capabilities was published in the news report on news.bpdata.com, where cutting development and operating costs takes center stage.

The unit economics on the operating side look strikingly low: 100 verified alerts cost about 30 cents, a pace that means 1200 alerts per hour. One hundred reports cost 85 cents, illustrated with a report-every-10-minutes scenario. Summarizing one hour of video costs 23 cents, takes 10 minutes on a single GPU, and yields six summaries per hour. Roughly a thousand visual search queries per hour translate to 3.37 dollars. Flexibility is presented as a feature: GPU budget can serve operator interaction by day and generative jobs like reports and summaries by night, with one hour of cloud GPU spend claimed to run all these capabilities together.

The engine of the efficiency gain is introduced as adaptive sampling and explained intuitively. The classical approach converts every patch of a 1080p frame into tokens for the model, while the adaptive system watches temporal and spatial change. Only tokens from moments and regions where change concentrates travel to the large model; the rest is pruned. The claimed result is bold: latency drops by 60 percent while the same vision-language model serves 40 percent more concurrent streams. Efficient sampling, previously shipped with fixed pruning rates, becomes an adaptive per-patch per-frame decision in version 3.3. On the Cosmos 3 streaming NIM side, a rolling window processes every frame in sequence, pushing the alert target below 300 milliseconds.

The measured capacity figures also clarify the hardware ladder: a single H100 handles 60 concurrent streams feeding into the vision-language model. Scenarios that skip feeding every frame into the model fit comfortably on far smaller edge hardware such as AGX or DGX Spark. The new support list also names Jetson Orin, Orin NX and GB300. The same architecture therefore scales from the datacenter to the factory edge, with hardware chosen by stream count, resolution and model preference. Sharing one skill set between a light edge setup and a heavy central one strengthens the claim of going everywhere with a single architecture.

Getting started and limits

Developer profiles keep the system modular with four validated starting points: base, alerts, long-video summarization and search. The skill starts from the closest profile to the request and applies only the needed delta, deduplicating shared services. When two capabilities need the same detector they share one detector, while Kafka and Elasticsearch are shared with separate indices where required. All 3.3 features can be tried today from the main branch of the GitHub repository and the nightly builds in the container registry. This open development line and nightly build rhythm run in the VSS repository on github.com, with the official 3.3 release tag planned to land with GTC Berlin.

Three stops are recommended for getting started: the cloud demo page, the GitHub repository and the documentation. One-click cloud deployment offers trial capacity on two RTX Pro cards, and Docker Compose plus Helm charts can be modified to fit. Kubernetes charts, security scans and production-ready NIM packages stand by for the production transition. Use cases seem endless: industrial inspection, warehouse and logistics, sports analytics, smart spaces and procedure validation stand out. A fifty-camera warehouse is said to connect through existing RTSP infrastructure and video management systems untouched, with hardware sized by resolution and processing frequency.

The zero data retention question gets a direct answer on privacy and compliance. Out of the box the reference architecture ships with a video management system and Elasticsearch, with default retention set to 24 hours in the index. Yet the system can also be installed with no video management and no database, and retention policies can be changed on both layers. On behavior analytics a sample microservice is provided: metadata flows into the message bus, heuristics such as overlapping boxes and velocity calculations run there, and anyone can attach their own service to the same bus. Combining thermal and optical cameras is considered feasible: a separate model that understands thermal input correlates raw temperature data with visual information.

Visualization: nodesdaily AI

Key moments

  1. Opening: the single-prompt vision agent claim
  2. VSS 3.3 updates: skill and Adaptive EVS
  3. Juice factory scenario setup
  4. Architecture proposal and deploy approval
  5. Fill analysis extension and masks
  6. Bottle 8105 questioned in chat
  7. Cost math: build and alert spend
  8. Q&A and getting-started steps

AI commentary

"My stance as narrator is clear: the single-prompt setup claim is backed by a real on-stage demo, but the cost and speed figures were measured only on NVIDIA hardware. Any developer reading this should measure with their own cameras in a trial deployment before making a production decision."

AI assessment

The strongest counterargument concerns generalization: a single bottling line demo, however polished, does not prove the same pipeline will behave on a different factory floor with different lighting, angles and bottle shapes. Vision-language verification reduces false alarms but it is not infallible, and the stream never shows a missed event or a wrong explanation being corrected. Readers should treat the accuracy impression as a demonstration, not a benchmark, until independent measurements on their own footage exist.

Several important details remain missing for a production decision. Latency and throughput figures come only from the presenters with no independent reproduction, GPU memory requirements per profile are barely discussed, and the calibration workload for a real multi-camera plant is mentioned only in passing. The zero-retention discussion stays at the level of configuration hints rather than a tested compliance recipe. A serious evaluator would want sizing tables, failure rates and retention proofs before signing off.

The speakers interest is openly commercial: one leads product management for the VSS blueprint and the other works in Metropolis technical marketing, so every number serves a launch narrative ahead of GTC Berlin. That does not make the figures false, but it explains which questions get airtime and which do not. Costs are computed on cloud pricing for a single RTX Pro card and assume a favorable workload mix. Treat the dollar amounts as directional, and re-run them against your own query patterns and camera counts.

The practical takeaway is a staged path: start from the nightly builds on github.com, try the one-click cloud deployment with your own recorded footage, and measure alerts per dollar on two or three cameras first. Only then decide between edge boxes and datacenter cards, and settle retention and calibration procedures before connecting fifty streams. The build.nvidia.com demo page is the fastest way to see what a finished agent feels like, while the official docs on docs.nvidia.com answer the configuration questions that follow.

Sources

7 links; no other published story cites them. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

nvidia · vss 3.3 · cosmos · artificial intelligence · video analytics · metropolis

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review permitted source passages, versions and origins.

KAYNAKLARLA OKU

Bu haberi açalım.

Hesap kontrol ediliyor…