Back to feed

From Laptop to Cluster: Taking an AI Agent Live with vLLM on OpenShift

A Red Hat walkthrough moves a three-tool agent off a laptop — where Qwen3 reasons through Ollama for free — onto an OpenShift cluster served by vLLM, changing nothing in the agent and everything around it.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — 6MuAOFfJk2w
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

An agent running on a laptop is a dream for exactly one person: Qwen3 does the reasoning, Ollama does the serving, no message data ever leaves the machine, and the bill is zero. But the dream ends the moment the lid closes for the commute home, or a second person wants to tap the same mind; a one-machine, one-user rig collapses under multi-user reality. The presenter builds the whole question from that fracture point: how do we turn this quiet desktop setup into a service a crowd can use, without falling into the one-machine trap ?

The starting rig is lean: a three-tool agent carrying a calculator, a web searcher, and a GitHub MCP connection, with reasoning supplied by the Qwen3 open-weight model and serving handled by Ollama on the laptop. Because message data never leaves the box, there is no privacy headache and no cost line; trial-and-error freedom is total. The host frames this as one-person, one-machine comfort, then states the production question crisply: how does the same agent live on without depending on a laptop with its lid shut?

The talk's strongest idea lands in the first minute: nothing inside the agent changes at all — what changes is where the model gets served and where the agent runs. There are two moves: first Ollama gives way to an inference server tuned for concurrent traffic, then the whole rig is containerized and moved onto an OpenShift cluster. The same shape works with outside endpoints such as Claude or GPT; the agent does not care where reasoning comes from as long as it has an endpoint abstraction to talk to. The Unsloth docs describe servers like Ollama and vLLM connecting through the same OpenAI-compatible interface pattern, which is why the laptop-to-cluster difference shrinks to three lines in a .env file.

First move: the model layer

Move one is the model layer. The laptop's Ollama setup is replaced by vLLM serving through OpenShift AI; vLLM is a Berkeley-born open project that, according to the RedHat blog introduction, was built for accelerator-memory efficiency and speed, gathering more than 40,000 stars on GitHub. As TheNewStack's internals piece explains, the server carries stacked calls on one model efficiently through continuous batching and paged-attention techniques, supporting families from Llama and Mistral to Granite and DeepSeek. In the AI Hub flow the model is pulled from a HuggingFace URI address, accelerator resources and a runtime are picked, and the deployment produces an internal link for in-cluster access, an external link for outside access, plus an API key that authenticates requests. The .env diff is three lines: the cluster's vLLM service instead of localhost Ollama, the fresh key, and the model identity vLLM expects.

The mind in the machine is Qwen3. Per the official announcement published on GitHub, the flagship 235B-A22B trades blows with top-tier models like DeepSeek-R1 and Gemini-2.5-Pro on coding and math, while a small mixture-of-experts build such as 30B-A3B and dense builds from 0.6B to 32B ship as open weights under Apache 2.0. A long context window and a hybrid thinking mode make the family a natural fit for tool-calling agent workloads; the free laptop rig in the demo is the everyday face of that licensing openness.

Second move: container and cluster

Move two is the agent itself. First a login command is grabbed from the OpenShift console and the cluster is entered; then the image is built with Podman, shipped through Quay, and rolled out — Docker does the same job. Make commands in the repo compress the build, push, and deploy steps into single keystrokes. A minute later the model endpoint and the agent run as side-by-side pods in the same cluster; the pod is scheduled, the image is pulled, and the agent that lived on a laptop a minute ago now lives as a pod on OpenShift. Not a line of code changed — only the home address did.

In the live trial the agent is first probed with a web-search and an arithmetic question, then asked to read an issue on GitHub and open a new one, with the same MCP tool shouldering reads and writes from inside the cluster. Chat logs are watched in the RedHat OpenShift AI playground instead of a local terminal; everything from the model to the agent to the playground runs in the cluster, and the setup shifts to in-cluster log watching.

Deployed is not production-ready

The finale draws an honest boundary: the pod can read any file inside its own container and reach out to nearly anything it wants; behavior visibility, resource limits, safe multi-team use, permissions assigned to the agent, and behavior scoping are all absent. Under the zero trust lens of the CNCF piece, unrestricted network traffic and open endpoints in AI workloads on clusters are the expensive edition of classic problems — a stolen key is not a small leak but a swelling bill. The open Agent Sandbox design covered by InfoQ builds isolation with a gVisor curtain, while the ephemeral-sandbox direction Google presented at KubeCon targets the same gap. The repo and the manifest wait in the links; the lockdown work is left as the next video's subject.

Visualization: nodesdaily AI
MoveGist
Model layerOllama out, vLLM in on OpenShift AI
Agent layerPodman built it, Quay shipped it, pod runs it
Missing pieceVisibility, quotas, permissions come later

Key moments

  1. Laptop rig and its three tools
  2. Model rollout on vLLM
  3. Image build and cluster rollout
  4. Live trial and safety warning

AI commentary

"The demo's framing is admirably honest: deployment is a change of stage, not a rewrite. But the security ledger is left for a sequel, and readers should pencil visibility and permissions into day one."

AI assessment

The strongest counter-argument comes from cost and scale: the vLLM swap is not free, demanding accelerator-memory planning, driver fit, and a permanently running cluster bill. For a two-user internal trial, Ollama's simplicity is more than enough; moving to a cluster at small load is driving a tank to the corner shop. With no concurrency figures or latency measurements in the demo, the scale threshold stays foggy: at how many simultaneous requests should the move happen — the talk never says.

The missing list is not short: which accelerator profile carries how many requests, how far latency falls, and what the monthly bill looks like are all absent. The security chapter is deferred to a promised sequel, so the word production arrives early, before permissions, quotas, and network policy are discussed. Viewers should watch this as a moving rehearsal rather than a setup recipe; whoever ships a quota-free, permission-free pod walks through a storm with an umbrella.

The presenter's position deserves a note too: the story airs on a Red Hat channel, with OpenShift AI and Quay in the natural lead roles, and no side-by-side cost comparison against managed outside endpoints. This content should be watched through a product-tour filter; vLLM's virtues are real, yet the stage belongs to the host. No budget decision should rest on comparison figures until independent measurements confirm them.

The portable lesson survives all that: abstract the reasoning endpoint into an address, key, and model-identity triple, and write the agent to be endpoint-independent so the endpoint can change tomorrow while the agent lives on. Start small, but never push visibility and permissions to day two; never shelve sandbox and quota work as a sequel topic. The free laptop rig is production's rehearsal — and rehearsals that are taken seriously get to move.

Sources

8 links; no other published story cites them. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

artificial intelligence · vllm · openshift · qwen3 · kubernetes · mcp

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review source passages, versions and origins.

READ WITH SOURCES

Understand this story.

Checking your account…