Build log · active

Aftershock build log

The build journey of Aftershock — a disaster-response society of Qwen agents built for the Qwen Cloud Global AI Hackathon. The claims that survived scrutiny: written doctrine lifts protocol conformance (credible at p=0.031; 95% on the NYC-Ida demo run), and six small models out-deliver one big one at ~65% better lives-per-dollar — alongside the negative results, and the lives headline we firmed and walked back, that got us to the honest numbers.

Qwen Cloud Agent society Honest benchmarks

Dispatches

11 LOGGED
LOG-011 8 MIN
12 models, 10 familiesNo solo beats the cheap societyFrontier ties on lives (a walk-back)Society wins ~12x on cost

We only ever tested Qwen. So we ran the headline against eleven other models — and had to correct how we say it

A reader's fair critique: every result in this log uses Qwen, so maybe 'a cheap coordinated society beats one big model' is a Qwen artifact. I built a family-agnostic provider (one OpenRouter endpoint; Qwen-only request fields stay off the non-Qwen hosts so the DashScope path is byte-identical) and ran the solo arm on twelve models from ten families. Verdict: no model's solo BEATS the cheap all-flash Qwen society on lives — but the eight frontier models (GPT-5, Gemini 3.1 Pro, Claude Opus 4.8, Grok, DeepSeek V4 Pro/Flash, Kimi K2.7, GLM 5.2) TIE it, every paired sign test p>=0.29, at 3-14x the cost. That's an honest correction to how I'd been phrasing it: on Qwen-only data a big Qwen solo sat at the swarm floor, so 'coordination beats a big model' read like a lives claim; cross-family a frontier solo reaches the outcome ceiling, and the society's win is on cost-efficiency (~12x on lives-per-dollar), holding across ten families. The honest dent: DeepSeek V4 Flash ties on lives at 4x better cost. And a clean cross-family capability floor: below the frontier solos fall off and an 8B model collapses. Spend ~$14.5.

Read dispatch
LOG-010 9 MIN
10× model spend → +0 liveslives-per-$ collapses 9.6×Coordination is friction-gatedAt abundance, coordination backfires (p=0.008)

We paid 10× for a bigger model and saved zero extra lives — then watched coordination backfire when nothing was scarce

I borrowed two open questions about agent societies and answered them on Aftershock, whose conserved lives-saved metric lets you separate 'the agent is good' from 'the substrate carried it.' (1) Does a bigger model decide better? Swapping the whole society roster flash → plus → max over 10 paired seeds: lives flat (106/107.5/107, all sign tests p>0.45), cost 9.7×, lives-per-dollar collapses 9.6×. Above the capability floor, model size buys nothing — a GPU-capex KILL; the cheap discipline lever (doctrine, +0.125 conformance at $0) moves what model scaling doesn't. (2) When does coordination matter? Sweeping pools from abundance to famine, the society-vs-swarm edge is an inverted-U: at abundance the auction is pure overhead and the swarm wins (p=0.008, the strongest cell in the study); it peaks at moderate scarcity (PoA 1.18×, suggestive p=0.289); it disappears at the floor where everything fails. scripted ≈ society at every level → the lever is coordination structure under contention, not model quality. Honest about what stays suggestive at n=8.

Read dispatch
LOG-009 11 MIN
Price of anarchy 1.11xCoordinated > uncoordinatedsolo ≈ swarmgap still suggestive (p=0.118)

We put a number on coordination: the price of anarchy in an agent society

A reader asked if the society-beats-swarm theme connects to Nash equilibrium. It does, cleanly: Aftershock is a common-pool resource game (finite shared rescue units, negative externalities), the protocol-free swarm is the uncoordinated price-of-anarchy baseline, and the society's per-tick auction + written doctrine is a coordination mechanism plus a correlation device. I built `aftershock poa` to measure it as a bounded efficiency — fraction of imperiled lives saved, from the sim's exact accounting. Result: both coordinated arms (society 67.3%, scripted central heuristic 66.2%) beat both uncoordinated ones (solo 58.9%, swarm 58.1%), and the big solo model sits at the swarm's anarchy level — coordination beats raw model size here. The pairwise society-vs-swarm gap stays suggestive (+6.7 efficiency points at n=15, p=0.118). I'm honest about what we can't claim (the agents aren't equilibrium-solvers; the true optimum is intractable; on the brutal real NYC-Ida pack the order flips). Then: the experiments that would firm it (a self-enforcement test, a central-planner oracle, auction strategyproofness) and how to turn the lens on real incident-response dispatch — with verified citations.

Read dispatch
LOG-008 7 MIN
Lead with what survivesConformance credible · p=0.031+8.9 lives stays 'suggestive'Observatory shows the evidence now

We firmed our headline until it broke. Here's the one that didn't.

Our marquee number didn't survive being firmed (Log 007: +28 → a suggestive +8.9, p=0.118). That forced an honest question: if your flashiest result is only suggestive, what do you headline instead? Answer — the claims a skeptic can't knock down. Written doctrine lifts protocol conformance credibly (+0.125, n=6, p=0.031, positive on all 11 seeds, zero lives cost; 95% on the NYC-Ida demo run), and six cheap Qwen models out-deliver one big model at ~65% better lives-per-dollar while matching hand-tuned expert heuristics on lives. The +8.9 lives edge stays in the story, labeled suggestive — not buried, not inflated. And the observatory, which had been greeting judges with 'No runs found,' now loads that evidence on landing.

Read dispatch
LOG-007 8 MIN
+28 → +8.9 (n=15)Doctrine now credibleTracker caught itselfFirm your headline

The protocol was 'worth 28 lives.' At fifteen seeds it's worth a caveat.

I had hackathon credits and a backlog of experiments. Instead of burning them, I built a thin experiment tracker — a provenance stamp on every result plus one queryable index — and on its first real use it labeled a credible conformance result as 'noise' because it was reading the wrong metric, so I fixed it. Then it firmed a real win (doctrine conformance, now credible at p=0.031) and demoted our marquee one: the '+28 lives' society-vs-swarm headline, firmed from five paired seeds to fifteen, collapsed to a suggestive +8.9 lives — the CI excludes 0 but the sign test doesn't clear significance — one seed had carried the original. We rewrote the README, the submission, and the evidence pack to the honest figure. The cost trim's conformance dip turned out small-but-real. Firm your proudest number first; the tracker earns its keep by demoting you.

Read dispatch
LOG-006 8 MIN
Wrong p-valueCherry-pick caught42 / 47 tracedAudit your honesty

We built a proof pack so judges could check our numbers. Auditing it ourselves found a wrong p-value, a fabricated source, and a cherry-picked run.

Near the deadline I stopped adding features and built the proof layer instead — a Decision Receipt for any ruling, confidence intervals and significance on the bench, and a one-page Evidence Pack that ties every headline number to a source file. Then a multi-agent adversarial pass re-derived every figure from source. 42 of 47 traced exactly; five did not — including a swarm p-value cross-contaminated from another row (0.375 vs 0.0625), a pack_digest cited as a JSON field that doesn't exist, and a 'flagship' NYC-Ida run shown at 0.95 conformance with no mention that it saved 8 of 90 lives. The honesty surfaces I built to impress judges were all overclaiming until I audited them. Intending to be honest isn't enough; you have to check your own receipts.

Read dispatch
LOG-005 9 MIN
Outcome-neutral trap−14% costCapability floorConformance

The fix that would have only fooled the scoreboard — and the one tuning that actually paid.

I used the harness to tune the society across four levers. Doctrine buys conformance at no lives cost (resolving an old n=1 scare). The infra agent's worst rule is a model-capability floor, not a prompt bug — so it shipped as an opt-in mode, not the default. And I nearly built a 'guard' whose only job was to game a metric the system already handles for free. The one unambiguous win: trimming the re-sent prompt cut cost 14% for +21% lives-per-dollar.

Read dispatch
LOG-004 8 MIN
Measure firstNegative resultFalse positive caughtStatistics

Build the ruler first. It killed our biggest feature — and a +16-life win that wasn't real.

I built the measurement harness before tuning the agents. It killed the backlog's biggest planned lever (a pathology that doesn't occur), caught a +16-life society-vs-swarm 'win' that collapsed to noise at 11 seeds, and forced an honest caveat on our own +28 headline. Three 'don'ts' worth more than a feature.

Read dispatch
LOG-003 6 MIN
ObservabilityCaught in reviewFrontendHonesty

We drew the agent auction on the map. A review caught it pointing at the wrong district.

We rebuilt the observatory map into a Mission Control view that draws the agent society's resource auction live — contested districts linked, loser to winner. Then a code review caught it pointing each loss at the wrong winning district. Making the mechanism visible meant making the picture honest.

Read dispatch
LOG-002 6 MIN
Negative resultAblationCost · latencyMulti-agent

We added native function calling. The benchmark told us to turn it off.

An honest ablation: native Qwen Cloud function calling, benchmarked on identical seeded worlds, cost ~2× for no lives benefit. Why per-call tool-schema overhead dominates in high-frequency multi-agent systems — and why JSON contracts stayed the default.

Read dispatch
LOG-001 9 MIN
ArchitectureBenchmarkFoundations

When does a society of small Qwen models beat one big model? Building Aftershock.

The founding question and the architecture behind it — splitting a disaster response across six specialized agents that negotiate scarce resources, and the seeded-world harness that decides whether the society actually pays for itself.

Read dispatch