Aftershock build log
The build journey of Aftershock — a disaster-response society of Qwen agents built for the Qwen Cloud Global AI Hackathon. The claims that survived scrutiny: written doctrine lifts protocol conformance (credible at p=0.031; 95% on the NYC-Ida demo run), and six small models out-deliver one big one at ~65% better lives-per-dollar — alongside the negative results, and the lives headline we firmed and walked back, that got us to the honest numbers.
Dispatches
11 LOGGEDWe only ever tested Qwen. So we ran the headline against eleven other models — and had to correct how we say it
A reader's fair critique: every result in this log uses Qwen, so maybe 'a cheap coordinated society beats one big model' is a Qwen artifact. I built a family-agnostic provider (one OpenRouter endpoint; Qwen-only request fields stay off the non-Qwen hosts so the DashScope path is byte-identical) and ran the solo arm on twelve models from ten families. Verdict: no model's solo BEATS the cheap all-flash Qwen society on lives — but the eight frontier models (GPT-5, Gemini 3.1 Pro, Claude Opus 4.8, Grok, DeepSeek V4 Pro/Flash, Kimi K2.7, GLM 5.2) TIE it, every paired sign test p>=0.29, at 3-14x the cost. That's an honest correction to how I'd been phrasing it: on Qwen-only data a big Qwen solo sat at the swarm floor, so 'coordination beats a big model' read like a lives claim; cross-family a frontier solo reaches the outcome ceiling, and the society's win is on cost-efficiency (~12x on lives-per-dollar), holding across ten families. The honest dent: DeepSeek V4 Flash ties on lives at 4x better cost. And a clean cross-family capability floor: below the frontier solos fall off and an 8B model collapses. Spend ~$14.5.
Read dispatch →We paid 10× for a bigger model and saved zero extra lives — then watched coordination backfire when nothing was scarce
I borrowed two open questions about agent societies and answered them on Aftershock, whose conserved lives-saved metric lets you separate 'the agent is good' from 'the substrate carried it.' (1) Does a bigger model decide better? Swapping the whole society roster flash → plus → max over 10 paired seeds: lives flat (106/107.5/107, all sign tests p>0.45), cost 9.7×, lives-per-dollar collapses 9.6×. Above the capability floor, model size buys nothing — a GPU-capex KILL; the cheap discipline lever (doctrine, +0.125 conformance at $0) moves what model scaling doesn't. (2) When does coordination matter? Sweeping pools from abundance to famine, the society-vs-swarm edge is an inverted-U: at abundance the auction is pure overhead and the swarm wins (p=0.008, the strongest cell in the study); it peaks at moderate scarcity (PoA 1.18×, suggestive p=0.289); it disappears at the floor where everything fails. scripted ≈ society at every level → the lever is coordination structure under contention, not model quality. Honest about what stays suggestive at n=8.
Read dispatch →We put a number on coordination: the price of anarchy in an agent society
A reader asked if the society-beats-swarm theme connects to Nash equilibrium. It does, cleanly: Aftershock is a common-pool resource game (finite shared rescue units, negative externalities), the protocol-free swarm is the uncoordinated price-of-anarchy baseline, and the society's per-tick auction + written doctrine is a coordination mechanism plus a correlation device. I built `aftershock poa` to measure it as a bounded efficiency — fraction of imperiled lives saved, from the sim's exact accounting. Result: both coordinated arms (society 67.3%, scripted central heuristic 66.2%) beat both uncoordinated ones (solo 58.9%, swarm 58.1%), and the big solo model sits at the swarm's anarchy level — coordination beats raw model size here. The pairwise society-vs-swarm gap stays suggestive (+6.7 efficiency points at n=15, p=0.118). I'm honest about what we can't claim (the agents aren't equilibrium-solvers; the true optimum is intractable; on the brutal real NYC-Ida pack the order flips). Then: the experiments that would firm it (a self-enforcement test, a central-planner oracle, auction strategyproofness) and how to turn the lens on real incident-response dispatch — with verified citations.
Read dispatch →We firmed our headline until it broke. Here's the one that didn't.
Our marquee number didn't survive being firmed (Log 007: +28 → a suggestive +8.9, p=0.118). That forced an honest question: if your flashiest result is only suggestive, what do you headline instead? Answer — the claims a skeptic can't knock down. Written doctrine lifts protocol conformance credibly (+0.125, n=6, p=0.031, positive on all 11 seeds, zero lives cost; 95% on the NYC-Ida demo run), and six cheap Qwen models out-deliver one big model at ~65% better lives-per-dollar while matching hand-tuned expert heuristics on lives. The +8.9 lives edge stays in the story, labeled suggestive — not buried, not inflated. And the observatory, which had been greeting judges with 'No runs found,' now loads that evidence on landing.
Read dispatch →The protocol was 'worth 28 lives.' At fifteen seeds it's worth a caveat.
I had hackathon credits and a backlog of experiments. Instead of burning them, I built a thin experiment tracker — a provenance stamp on every result plus one queryable index — and on its first real use it labeled a credible conformance result as 'noise' because it was reading the wrong metric, so I fixed it. Then it firmed a real win (doctrine conformance, now credible at p=0.031) and demoted our marquee one: the '+28 lives' society-vs-swarm headline, firmed from five paired seeds to fifteen, collapsed to a suggestive +8.9 lives — the CI excludes 0 but the sign test doesn't clear significance — one seed had carried the original. We rewrote the README, the submission, and the evidence pack to the honest figure. The cost trim's conformance dip turned out small-but-real. Firm your proudest number first; the tracker earns its keep by demoting you.
Read dispatch →We built a proof pack so judges could check our numbers. Auditing it ourselves found a wrong p-value, a fabricated source, and a cherry-picked run.
Near the deadline I stopped adding features and built the proof layer instead — a Decision Receipt for any ruling, confidence intervals and significance on the bench, and a one-page Evidence Pack that ties every headline number to a source file. Then a multi-agent adversarial pass re-derived every figure from source. 42 of 47 traced exactly; five did not — including a swarm p-value cross-contaminated from another row (0.375 vs 0.0625), a pack_digest cited as a JSON field that doesn't exist, and a 'flagship' NYC-Ida run shown at 0.95 conformance with no mention that it saved 8 of 90 lives. The honesty surfaces I built to impress judges were all overclaiming until I audited them. Intending to be honest isn't enough; you have to check your own receipts.
Read dispatch →The fix that would have only fooled the scoreboard — and the one tuning that actually paid.
I used the harness to tune the society across four levers. Doctrine buys conformance at no lives cost (resolving an old n=1 scare). The infra agent's worst rule is a model-capability floor, not a prompt bug — so it shipped as an opt-in mode, not the default. And I nearly built a 'guard' whose only job was to game a metric the system already handles for free. The one unambiguous win: trimming the re-sent prompt cut cost 14% for +21% lives-per-dollar.
Read dispatch →Build the ruler first. It killed our biggest feature — and a +16-life win that wasn't real.
I built the measurement harness before tuning the agents. It killed the backlog's biggest planned lever (a pathology that doesn't occur), caught a +16-life society-vs-swarm 'win' that collapsed to noise at 11 seeds, and forced an honest caveat on our own +28 headline. Three 'don'ts' worth more than a feature.
Read dispatch →We drew the agent auction on the map. A review caught it pointing at the wrong district.
We rebuilt the observatory map into a Mission Control view that draws the agent society's resource auction live — contested districts linked, loser to winner. Then a code review caught it pointing each loss at the wrong winning district. Making the mechanism visible meant making the picture honest.
Read dispatch →We added native function calling. The benchmark told us to turn it off.
An honest ablation: native Qwen Cloud function calling, benchmarked on identical seeded worlds, cost ~2× for no lives benefit. Why per-call tool-schema overhead dominates in high-frequency multi-agent systems — and why JSON contracts stayed the default.
Read dispatch →When does a society of small Qwen models beat one big model? Building Aftershock.
The founding question and the architecture behind it — splitting a disaster response across six specialized agents that negotiate scarce resources, and the seeded-world harness that decides whether the society actually pays for itself.
Read dispatch →