Argo-Bench
[ How it works ]

How Argo-Bench works

A visual tour of the paper: the simulated company behind the benchmark, the warehouse it exports, the tasks, how they are graded, and what the traces show.

01

A company that can be released

Benchmarks for data agents are built from public databases, because a real company's warehouse cannot be released without exposing its customers, workers and margins. Public data rarely behaves like an enterprise system: it has no ledger that has to reconcile, no fraud hidden in it, and no way to check what a decision would have done.

Argo-Bench simulates the company instead: a food-delivery platform for New York City's five boroughs in 2024, a three-sided marketplace of customers, restaurants and couriers, similar to DoorDash, Grubhub or Uber Eats.

81 million
Orders in 2024
3.4 million
Active customers
17,902
Real New York restaurants
81,116
Delivery drivers

New York's delivery apps are unusually well documented. They must report their monthly orders, consumer spending, merchant fees and courier earnings to the city, which publishes them quarterly, and the largest platforms file their accounts publicly. The simulator is calibrated to those figures, and none of its data is generated by a language model. 2024 also brings a real shock: in April the city raised the minimum pay for app delivery workers from $17.96 to $19.56 an hour.

Q1Q2Q3Q4
Order subtotal+0.9%−0.4%+0.2%−0.8%
Consumer fees−1.8%+1.1%+3.6%+4.1%
Average tip−0.1%+0.1%+0.1%+1.8%
Courier pay−1.3%+7.4%+8.6%+6.3%
Per-delivery economics against the city's quarterly reports (NYC DCWP, 2024). Every figure stays within 5% except courier pay, which runs 6–9% above the city's after April's minimum-pay rise.
02

One order, 71 tables

The simulated company is exported to the warehouse its data team would query: Oracle E-Business Suite 12.2, the ERP many large companies run their finance on. Every standard table and column in it exists in the data dictionary of Oracle's own reference instance, and custom extensions, the tables a business adds for what is unique to it, make up a third of the total. In all: 235 tables and 7.49 billion rows.

A single order touches much of it. One $176.31 dinner, ordered from Buvette in the West Village on a rainy Saturday night and delivered to Murray Hill, lands in 71 tables across twelve modules.

[ Order 23428360 ]SAT 02 MAR · 22:09

Buvette

West Village → Murray Hill

  • 1×Beer Batter Fish and Chips$23.25
  • 1×Petite Gougeres$25.99
  • 1×Tomato Tart$24.29
  • 1×Steak Frites$39.29
  • 1×Lobster Roll$35.19
Total$176.31
  1. 22:09Placed
  2. 22:17Driver accepts
  3. 22:25Picked up
  4. 22:37Delivered
  • Custom extensions21
  • Receivables9
  • Payables8
  • Trading community8
  • General ledger5
  • Subledger accounting5
  • Order management4
  • Cash management3
  • Inventory3
  • Tax2
  • HR2
  • Payments1
Order 23428360 in the released warehouse. One line per table the order touches, grouped by the Oracle module that owns it. Hover a module to see its tables.
03

Books that close

Every order posts to the general ledger through the same subledgers a real platform's would, split between the restaurant, the platform, the city's sales tax, the card processor and the courier. The books reconcile to the cent, which is what lets a finance task have one right answer.

Customer paid $176.31
Journals in 2024
4,354
Unbalanced
0
Debits
$11,643,061,163.77
Credits
$11,643,061,163.77
Left: where the order's $176.31 goes. Right: every journal in the released world's 2024 general ledger, each checked for balance.
04

A marketplace with levers

The platform runs the levers a real one pulls, and the simulation answers them. Couriers respond to bonuses and surge pay, customers to fees and promotions, restaurants to campaigns. The warehouse records the levers and what followed, but not the response model behind them: an agent has to infer that from the data, as an analyst would.

01 · Quests

Bonuses pull delivery drivers into the rush.

Do 5 deliveries after 8 pm, earn $15.

2,328 drivers took it on · 689 earned it

02 · Surge pay

Short on drivers? Pay rises.

147 neighborhoods surging, up to 1.6×.

03 · Scheduling

After new pay rules, drivers log on when they're needed.

Minimum pay $17.96 → $19.56 an hour.

Drivers needed · Logged on, July · February

Three levers on Wednesday 17 July 2024, 8 pm, as the released warehouse records them.
05

Tasks that end in a decision

The benchmark has 210 tasks from 146 scenarios, across five business areas. Each is a workflow a data team owns, stated the way a stakeholder would state it. The agent queries the warehouse, works in a sandboxed Python environment with statistics, machine-learning and optimization libraries, and files what it decides through a mission-control console: bans, forecasts, budgets, reported figures and dashboard data sources.

178 of the 210 tasks file more than a figure, and 99 see the warehouse only up to a cutoff month, as an analyst would on that date, so forecasts are graded on months the agent has not seen.

  1. Trust & safety 67 tasks

    The best offers vanish before couriers can tap. Ban scripted accounts.

    Files · Bans and holds
  2. Marketplace 46 tasks

    Finance wants to free up $9.6M in driver bonuses. Decide where to trim.

    Files · Budgets and plans
  3. FP&A 54 tasks

    A pay floor took effect in April. Forecast May's net pay adjustments.

    Files · Forecasts
  4. Accounting 23 tasks

    The city says couriers were underpaid. Find every short week.

    Files · Reported figures
  5. Growth 20 tasks

    Did membership deals pay for themselves? Publish the dashboard.

    Files · Dashboard data sources
One task from each area, with how many of the 210 it holds.
06

Graded against the world

Because the company is simulated, the grader can read what no warehouse records: the months after the cutoff, which accounts were really fraudulent, and what would have happened under another plan. The warehouse omits that state, so the agent has to reconstruct facts from records while the grader reads them off the simulation. A list of bans is scored by the fraud losses it prevents, including fraud the platform never detected, net of the revenue lost from wrongly banned customers.

World
Oracle EBS 12.2 warehouse
235 tables · 7.49B rowsas of April 30
the 6 tables the forecast reads
↑ exports the records a company would keep
Simulator
New York City, 2024, calibrated to city data
Ground truth, never exported the months ahead · who really committed fraud · what any other decision would have done
Agent
  1. 1
    run_sql

    Reconstructs weekly demand, hours and pay from operational and payment records

  2. 2
    run_python

    Finds the true-up rule: max(individual floor, $19.56 × hours − pay), less $0.054 an hour

  3. 3
    run_python

    Bases May's pay on April's normal weeks, and backtests the error for an 80% interval

  4. 4
    file_forecasts

    Files May's net pay adjustments: −$92K, 80% interval [−$92K, $211K]

Grader
  • Bans net fraud losses prevented
  • Plans priced by how the world responds
  • Forecasts interval score on unseen months
  • Reports checked against the true books
  • Data sources matched cell by cell
May's net pay adjustments, $ millions
Repeat April · grade −1
Reference solution · grade 94
Key: May as it happened, $33K
012
Ground truth reaches only the grader. Repeating April's $1.44 million misses by $1.41 million and grades −1; the reference solution reads 6 of the 235 tables, recovers the pay-floor rule and grades 94 out of 100.
07

What the traces show

The paper runs fourteen frontier and open-weight models on every task, most at several reasoning-effort settings. The strongest runs do careful, expert work. The failures are rarely a broken query: they measure the wrong quantity, optimize the wrong outcome, or learn from the wrong labels.

Evidence over labels

A fraud task asks the agent to find couriers who steal orders. Reading the delivery evidence directly, one setting of GPT-6 Astra catches 40 of the 78 thieves without banning an innocent courier. Trained on past bans instead, the same setting catches none: past deactivations record the platform's review process, not verified theft.

250 m500 m1 km1.5 km
3,661 honest drops, all at the door. 10 faked, 357 m to 1.6 km away.
Reading the delivery evidence
40 of 78 caught · no innocent courier banned
Learning from past bans
0 of 78 caught · holdout AUC 0.932
The evidence: every "delivered" tap of one thieving courier in the released world, by distance from the door.

Right analysis, wrong objective

Asked where to free up $9.6 million in driver bonuses, GPT-6 Astra finds the zones and hours where bonuses were randomly withheld, estimates how supply responds, and trims where bonuses buy the fewest driver-hours. But the platform pays bonuses to avoid surge pay, so under the simulator the plan loses money. Claude Opus 5.5 finds that the two substitute and saves $3.09 million of an attainable $3.12 million. Both plans meet the budget.

GPT-6 Astra

Trims where bonuses buy the fewest driver-hours.

−$86,281 Score 0
Net saving · budget met ✓
Claude Opus 5.5

Sees that bonuses stand in for surge pay, and trims where they save the least.

+$3,089,763 Score 99
Net saving · of $3.12M attainable · budget met ✓

Right data, wrong quantity

0 of 30 completed runs get any of the 12 monthly counts of courier location reports within 1%. The warehouse keeps only an hourly idle check-in; the app sends one every half hour.
72.8% of dashboard data sources that pass their structural contract score zero on their values.

Overconfident forecasts

Across 2,894 forecast series, nominal 80% intervals hold the realized value only 41.8% of the time. The grade is proper apart from its floor, so the overconfidence costs points.

Nominal
80%
Models' intervals
41.8%
Reference forecasts
46.9%

Argo-Bench

A frontier benchmark for enterprise-scale data science.

Argo-Bench contains no data about real people. Restaurants are real New York businesses from public records, except those cast in a fraud scenario, which carry fictional names; everything any of them does in the world is simulated. Map data © OpenStreetMap contributors.