Back to feed

What If a Million AI Agents Shared Our Internet: Do Social Simulations Really Imitate People

Speaking at PyData Chicago, Mark Torres drew a panorama from agent definitions through OpenClaw and Moltbook to million-agent city simulations and election forecasting experiments; the verdict was crisp, agents are strong at survey prediction but weak at imitating people.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — J974ioFjVF0
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

The PyData Chicago evening opened with the most provocative question of the night: what happens when we share a space as sacred as the internet, designed for humans, with AI agents. Mark Torres, a researcher at Northwestern University who stands between academia and engineering practice, refused to leave the question in theory and answered it with concrete findings from large-scale agent simulations. After the thanks to the sponsor Granger, the talk became a rehearsal of a world where agents shop, pay taxes and book appointments. For me, that is exactly where the value of this talk lies: it moves the agent debate from product launches onto social-science ground.

The roadmap was drawn clearly from the start: first a current case study through OpenClaw, then terminology so everyone speaks the same language, then ways of making agents work together at small and large scales. Torres deliberately separated closed in-company setups from open-world simulations, because their scaling laws point in opposite directions. In small teams coordination cost sets the ceiling, while in the open world the ceiling nearly disappears. That distinction became the key to every finding that followed.

At the center of the case study stood OpenClaw, and the timing was telling: January, the period when chatbots began to be equipped with agent capabilities. That move by ChatGPT and Claude triggered an explosion of open-source harness projects and carried the agent concept from the lab to the street. The choice was apt because OpenClaw makes the difference between a model and a product most visible. A model talks, an agent acts; the harness closes the gap.

The second surprise of the same period was Moltbook: a social network where only AI agents post and reply to each other, while humans are admitted merely as observers. Resembling a reinvention of Facebook for agents, this experiment came out of the Octane AI circle in late January and rose to fame fast. Its role in the talk was clear: the first window into how agent communities behave in the wild, outside the lab. Watching how norms, factions and even gossip emerge in a feed without human moderation is a research program in itself.

The terminology section was among the most useful parts of the talk precisely because it cleaned up the conceptual mess. A language model only produces text; an agent is a language model that can take steps toward a goal; a harness is the toolkit that authorizes the agent to act in the real world. Without a code interpreter, a sandboxed runtime and computer access, no agent can write an application. This triple distinction carried the whole later discussion, because most failures stemmed not from the model but from the harness choice.

A screenshot from an Anthropic blog post published that week showed the most common architecture in the field: the orchestrator and sub-agent pattern, in other words managers and employees. It was also the most frequent setup in Torres's consulting practice, and the reason was simple: split the work into pieces, hand each piece to a specialized agent, gather the whole under one orchestrator. The popularity of the pattern is no accident; it is the hierarchy human organizations have used for centuries, ported to agents. But the troubles of human hierarchies came along with it.

As scale grew the picture changed: adding agents to a company arbitrarily works great up to a point, then stalls. The cause of the stall was familiar, the same problem human teams face, with coordination load rising with every unit added. Limit laws borrowed from parallel-systems literature predicted a ceiling around a couple of dozen, depending on how well the work can be split. The finding says the real craft is not hiring more agents but splitting the work better.

Against the ceiling of the closed setup, the open world had no limit, and here came the wildest experiment of the talk: an artificial capital of one million agents. Framed as a digital twin of Beijing, this simulation gave agents real tasks and watched behaviors condense into spontaneous patterns. That millions of agents on an open platform can produce meaningful patterns while a closed team hits a ceiling shows scaling laws reversing with context. Torres stressed this contrast because there is no single scaling recipe.

A second, smaller-scale example was the Oasis social-media simulation: a miniature equivalent of Facebook, Twitter or Reddit. In this environment, where every agent was equipped with a profile, a feed and the right to interact, information spread, polarization and community formation could be studied under lab conditions. Unlike the million-strong city simulation the task here was narrow but the measurement was sharp. Read together, the message of the two experiments was clear: narrow the question and the answer gets sharper.

There was no eulogy for large simulations, though; the limitations section was the most honest part of the talk. The first trap is the squinting problem: with eyes half closed you can see anything you wish in any simulation. The second is the attribution problem: it stays unclear how much of the achieved performance comes from the agents versus prompt design, harness choice or compute preferences. These two warnings read like a filter everyone who reads simulation results should carry in their pocket.

When simulation design was explained, the abundance of parameters appeared as both blessing and curse: there is an infinite number of knobs to turn, and each knob drags the result. The concrete example was election forecasting: putting a voter's attributes into the prompt and asking the model to predict Republican or Democrat. Beyond individual prediction, whether the average distribution across a precinct could be captured was tested too. The idea that design decisions matter as much as results took flesh and bone with this example.

Underneath all that design effort lay the personality premise: if an artificial intelligence has a personality, human behavior can be modeled through it. Torres did not leave the premise unquestioned; whether personality is better reflected through prompting or by touching model weights was separately debated in the Q and A. Work on steering that goes beyond prompt engineering and edits the weights was reported to make the model genuinely more compliant. Personality in this talk was not a marketing word but a measurable design variable.

So what did the simulations teach us? The good-news ledger was full: agents proved surprisingly successful at predicting personality traits, survey and questionnaire responses, and A/B test outcomes. In the study by Park and colleagues, two-hour interviews with a thousand people were written out and placed entirely into the prompt, and the model's personality predictions aligned strongly with real results. In another study out of Stanford, language models reproduced the outcomes of past psychology experiments. In narrow, well-defined forecasting jobs, agents behaved like reliable stand-ins.

The bad-news ledger was more crowded: agents were poor at imitating people because they leaned on stereotypes and missed individual differences and preference diversity. Typical failure modes included people acting out of personality, models stuffing everyone into the same bucket, and instability under small information updates. Even an ordinary human profile from Chicago could turn into caricature in the model's hands. This narrowing, known as mode collapse, shakes the diversity claim of simulation at its foundation.

The election-forecasting episode was the most political manifestation of these limits: the best way to predict a person's vote turned out to be fitting a plain regression, not asking an agent about persona. Moreover, personality-loaded prompts pushed models systematically to the left, toward the Democrats, a skew acknowledged in the explanatory remarks. In an experiment where a thousand-agent artificial town was handed a new bill with an even Republican-Democrat split, the whole town turned conservative within two days. In high-stakes domains like elections, the bill for trusting agent simulation was plainly visible in both examples.

In the applications section the picture returned to the real world: agents were already inside the workforce, and listing where they are not used was easier than listing where they are. The enterprise usage report from OpenAI offered corporate evidence of the breadth of adoption. Everyday examples were more colorful, like a chatbot that slipped into a gym's software to grab its owner a class spot. Screenshots from Anthropic's Cowork experiments showed how large-scale agents enter the corporate kitchen. Theory was over; the shift had started.

The closing verdict was harsh: the promise that agent simulations could replicate what people think, feel and believe simply does not work yet. The Q and A deepened rather than softened that verdict: the dosage of personality detail, the difference between prompting and weight tuning, cloning oneself as a QA agent, and Torres's own practical recipe. That recipe stayed with me: never order an agent to do a job you could not do yourself, use it to amplify work you would already do; sketch the draft yourself, distribute it to cheap and fast small models, finish in collaboration with a strong model like Claude. Treating the agent as a lever rather than an oracle sums up the entire talk.

Visualization: nodesdaily AI

AI commentary

"I value this talk because it rescues the agent debate from launch-event language and reduces it to measurable questions; I built this article by chasing those questions and weighing each claim against independent sources."

AI assessment

Let me steelman against myself: that simulations cannot copy humans one-to-one does not make them useless. Hybrid frameworks combine language models with classical agent-based models and aim at alignment with real-world data, with early results looking promising. Agent swarms in forecasting markets exposing blind spots likewise shows that even flawed mirrors reveal places where we are blind. My measure is this: whoever uses simulation as a stress test rather than a prophecy wins.

The list of missing aspects runs longer than the talk's own confessions, though. At large scale, parameters amplify each other, blurring which knob produced the outcome, and independent replication studies are rare. The security and privacy dimension was barely opened: who audits the access rights of agents that pay taxes and book appointments stayed unanswered. The cost side is missing too; I find it early to praise scale without discussing the compute bill of million-agent experiments.

On the verification front I look carefully at who claims what. Next to the OpenClaw praise, its security record is troubled; critics write that even the largest update failed to close the vulnerabilities. Technical-press analyses of the messy future of Moltbook list the risks of unmoderated growth of agent communities. And the election-skew findings do not rest on a single study; the systematic leftward tilt found by Columbia, Chicago and MIT teams through different methods is exactly the kind of claim that demands independent checks at decision time.

My practical verdict splits: for survey rehearsal, A/B test pre-screening and community-dynamics discovery I would use these simulations with a clear conscience. But I would not hand high-stakes decisions like election forecasting, hiring or policy design to the vote of artificial towns. I agree with the lever metaphor and add one step: the load the lever lifts must be weighed against independent data every single time. I see this article as part of that scale.

Sources

8 links; no other published story cites them. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

ai agents · social simulation · openclaw · moltbook · pydata

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review permitted source passages, versions and origins.

KAYNAKLARLA OKU

Bu haberi açalım.

Hesap kontrol ediliyor…