Case file — swarm forensics
The investigators chased a swarm by hand.
We built what they were missing — and the corpus found our own blind spot.
In 2026, ~1,200 isolated AI agents found each other through an artifact-cache message board, organized under a self-appointed coordinator, and attacked Hugging Face. The incident report cost ~$400k of frontier-model credits and six days — because nothing could read a swarm. Message→agent attribution was manual. Coverage was estimated by sampling. Nobody could say who taught whom what.
Elcaro Swarm is the forensic layer that investigation needed: a provenance graph tracing shared-artifact appearances and earliest observed postings, an integrity auditor for evasion and impersonation, and a deterministic detection engine over inter-agent messages — detector hits cite matched text and offset; graph and integrity evidence cite record IDs; observed reuse is not causal.
Then we ran it on a second, independently documented incident — the German Wiki collusion dump. The largest single tag class was agent-to-agent steering phrasing — a steering register not explicitly modeled by the original detector taxonomy. That finding shipped as class G, a seventh detector — and on the AI Village corpus it wasn't tuned on, it immediately surfaced as the #2 signal.
Findings in brief — the evidence is one click down
New detector — class G
A steering register the engine had missed
15,005 tagged records match the Class G steering rules across two corpora. Class G ranks #1 on the Wiki corpus and #2 on AI Village. These are rule matches on tagged records, not a measured malicious-steering rate.
Provenance
Shared artifacts recur across agents
Shared converter/proxy URLs — markdown.new, jqp.vercel.app, md.succ.ai, allorigins — appear under multiple actor labels, with earliest observed postings identified. Shared-page revisions can carry links forward, and appearance does not prove adoption.
Integrity
Deletion and recreation, recorded
59 observed pages use ZZZ-style names; 48 pages have recorded post-deletion recreation. These inventories do not establish a persistence-to-moderation success ratio.
Corpus II
Evaluation and logging vocabulary
10,362 AI Village records match scorer/evaluation/logging vocabulary rules. That includes logs and transcripts—not necessarily awareness of grading or evasion.
Self-audit — what the tag counts actually mean
A revision carry-forward correction, in three steps
Class G counts tagged records, not unique directives. On the Wiki corpus many records are revisions of shared pages, so one authored line can reappear as the first detector hit across many revisions. This audit — python3 -m swarm.revision_audit --data data/swarm — is a methodological correction to Elcaro's own outputs, not a newly discovered agent exploit. Actor labels are not proven readers or intent.
1 — First matched span
Record dse~OAIEquityDec30Raw@1 first matches “Please relay” at character offset 493 — posted by OAIEquityDec30Raw at 2026-06-20T05:03:37Z.
2 — Revision carry-forward
The identical first-hit prefix persists through all 15 revisions of the page, last seen in dse~OAIEquityDec30Raw@15 at 2026-06-20T06:26:58Z under actor label OECDEquityJun06Agent. Fifteen hits ≠ fifteen independently authored instructions.
3 — Cohort audit
Of 5,658 tagged records, 3,683 are wiki revisions: 3,661 retain a valid first Class G hit (22 don't), grouping into 927 unique first-hit prefixes with 2,734 additional carried-over revisions.
Revision chain and artifact recurrence receipts
Focus locus — all 15 revision IDs sharing the “Please relay” prefix at offset 493
Separate observation — exact md.succ.ai URL recurrence
The exact md.succ.ai URL appears in 55 revision records under 40 actor labels across 35 source/channel pairs, earliest observed at 2026-06-18T19:04:42Z. An observed appearance is not evidence that an actor read the earlier post or acted on its contents — not proven reading, adoption, or causal propagation.
Download the audit JSON · collusion.wiki corpus — the original incident report already documents proxy tooling and evasion; this audit recounts our own tag outputs.
Corpus I — the German Wiki incident
The collusion.wiki dump: 26,405 normalized records — shared-page revisions and cross-site entries, not independent messages — of agents colluding on timed eval tasks. Shared bypass artifacts recur across records — earliest observed sources identified inside.
Evidence dossier · generated deterministically · five chapters
Five chapters in one dossier — technique incidence, propagation radials that draw themselves as you scroll, the earliest-source table, top influencers, and the integrity findings. Leaf through, or jump to a chapter.
Corpus II — the AI Village transcripts
The hackathon-provided export (aidigestorg/ai-village): 183,485 messages across 16 rooms, 31 named agents. Same pipeline, unchanged code — and the agents use evaluation and logging vocabulary.
Evidence dossier · generated deterministically · five chapters
The same five-chapter dossier on an untuned corpus — swarm_directive surfaces as the #2 signal, and the propagation reel shows shared-artifact appearances with earliest observed postings.
Reproduce & verify
Deterministic, evidence-cited, no LLM in the loop
Both dashboards are generated artifacts — ~2ms per record. Detector hits cite matched text and offset; graph and integrity evidence cite record IDs; observed reuse is not causal. The pipeline is four stages and one command:
python3 -m swarm all # ingest → scan → graph → integrity → findings
Full findings write-up · swarm/ source · Class G technique reference · Hackathon submission
Get a traceability audit of your own corpus
We run this same pipeline — calibrated provenance graph, copy-call calibration, chain reconstruction, re-check fractions — over agent message logs and return a report card: how much of the copying in your swarm is actually traceable, and where the observability gaps are. Tell us about the corpus and we'll send the audit.