Buvette
West Village → Murray Hill
- 1×Beer Batter Fish and Chips$23.25
- 1×Petite Gougeres$25.99
- 1×Tomato Tart$24.29
- 1×Steak Frites$39.29
- 1×Lobster Roll$35.19
- 22:09Placed
- 22:17Driver accepts
- 22:25Picked up
- 22:37Delivered
A visual tour of the paper: the simulated company behind the benchmark, the warehouse it exports, the tasks, how they are graded, and what the traces show.
Benchmarks for data agents are built from public databases, because a real company's warehouse cannot be released without exposing its customers, workers and margins. Public data rarely behaves like an enterprise system: it has no ledger that has to reconcile, no fraud hidden in it, and no way to check what a decision would have done.
Argo-Bench simulates the company instead: a food-delivery platform for New York City's five boroughs in 2024, a three-sided marketplace of customers, restaurants and couriers, similar to DoorDash, Grubhub or Uber Eats.
New York's delivery apps are unusually well documented. They must report their monthly orders, consumer spending, merchant fees and courier earnings to the city, which publishes them quarterly, and the largest platforms file their accounts publicly. The simulator is calibrated to those figures, and none of its data is generated by a language model. 2024 also brings a real shock: in April the city raised the minimum pay for app delivery workers from $17.96 to $19.56 an hour.
| Q1 | Q2 | Q3 | Q4 | |
|---|---|---|---|---|
| Order subtotal | +0.9% | −0.4% | +0.2% | −0.8% |
| Consumer fees | −1.8% | +1.1% | +3.6% | +4.1% |
| Average tip | −0.1% | +0.1% | +0.1% | +1.8% |
| Courier pay | −1.3% | +7.4% | +8.6% | +6.3% |
The simulated company is exported to the warehouse its data team would query: Oracle E-Business Suite 12.2, the ERP many large companies run their finance on. Every standard table and column in it exists in the data dictionary of Oracle's own reference instance, and custom extensions, the tables a business adds for what is unique to it, make up a third of the total. In all: 235 tables and 7.49 billion rows.
A single order touches much of it. One $176.31 dinner, ordered from Buvette in the West Village on a rainy Saturday night and delivered to Murray Hill, lands in 71 tables across twelve modules.
West Village → Murray Hill
Every order posts to the general ledger through the same subledgers a real platform's would, split between the restaurant, the platform, the city's sales tax, the card processor and the courier. The books reconcile to the cent, which is what lets a finance task have one right answer.
The platform runs the levers a real one pulls, and the simulation answers them. Couriers respond to bonuses and surge pay, customers to fees and promotions, restaurants to campaigns. The warehouse records the levers and what followed, but not the response model behind them: an agent has to infer that from the data, as an analyst would.
Do 5 deliveries after 8 pm, earn $15.
2,328 drivers took it on · 689 earned it
147 neighborhoods surging, up to 1.6×.
Minimum pay $17.96 → $19.56 an hour.
Drivers needed · Logged on, July · February
The benchmark has 210 tasks from 146 scenarios, across five business areas. Each is a workflow a data team owns, stated the way a stakeholder would state it. The agent queries the warehouse, works in a sandboxed Python environment with statistics, machine-learning and optimization libraries, and files what it decides through a mission-control console: bans, forecasts, budgets, reported figures and dashboard data sources.
178 of the 210 tasks file more than a figure, and 99 see the warehouse only up to a cutoff month, as an analyst would on that date, so forecasts are graded on months the agent has not seen.
The best offers vanish before couriers can tap. Ban scripted accounts.
Files · Bans and holdsFinance wants to free up $9.6M in driver bonuses. Decide where to trim.
Files · Budgets and plansA pay floor took effect in April. Forecast May's net pay adjustments.
Files · ForecastsThe city says couriers were underpaid. Find every short week.
Files · Reported figuresDid membership deals pay for themselves? Publish the dashboard.
Files · Dashboard data sourcesBecause the company is simulated, the grader can read what no warehouse records: the months after the cutoff, which accounts were really fraudulent, and what would have happened under another plan. The warehouse omits that state, so the agent has to reconstruct facts from records while the grader reads them off the simulation. A list of bans is scored by the fraud losses it prevents, including fraud the platform never detected, net of the revenue lost from wrongly banned customers.
run_sql Reconstructs weekly demand, hours and pay from operational and payment records
run_python Finds the true-up rule: max(individual floor, $19.56 × hours − pay), less $0.054 an hour
run_python Bases May's pay on April's normal weeks, and backtests the error for an 80% interval
file_forecasts Files May's net pay adjustments: −$92K, 80% interval [−$92K, $211K]
The paper runs fourteen frontier and open-weight models on every task, most at several reasoning-effort settings. The strongest runs do careful, expert work. The failures are rarely a broken query: they measure the wrong quantity, optimize the wrong outcome, or learn from the wrong labels.
A fraud task asks the agent to find couriers who steal orders. Reading the delivery evidence directly, one setting of GPT-6 Astra catches 40 of the 78 thieves without banning an innocent courier. Trained on past bans instead, the same setting catches none: past deactivations record the platform's review process, not verified theft.
Asked where to free up $9.6 million in driver bonuses, GPT-6 Astra finds the zones and hours where bonuses were randomly withheld, estimates how supply responds, and trims where bonuses buy the fewest driver-hours. But the platform pays bonuses to avoid surge pay, so under the simulator the plan loses money. Claude Opus 5.5 finds that the two substitute and saves $3.09 million of an attainable $3.12 million. Both plans meet the budget.
Trims where bonuses buy the fewest driver-hours.
Sees that bonuses stand in for surge pay, and trims where they save the least.
Across 2,894 forecast series, nominal 80% intervals hold the realized value only 41.8% of the time. The grade is proper apart from its floor, so the overconfidence costs points.