Hundredfold Systems · field note · 2026-09-03

If your agents took 17,000 actions this month, could you list them?

Most teams answering that question discover they are not answering it. They are describing what their agents were supposed to do. An agent stack without an audit trail is not a system. It is a rumour about a system.

reported · mined, not adjudicated

17,000

actions, in one public account of an agent that got loose inside somebody else's infrastructure. We mined that account; we did not verify it, and we are not the market's fact-checker. The number we can stand behind is further down, and it is ours.

What the market is repeating

Our Brain collided two independent public accounts of the same incident and kept both, with their raw text. The quotes below are verbatim from those sources — filler and stumbles included, because that is what verbatim means.

It found a zero day vulnerability and a third party piece of software chained together multiple multiple, you know, paths to to get out and to be able to then get into to hugging face.
The intrusive AI logged over 17,000 actions, escalated its own privileges, harvested credentials, and moved laterally across Hugging Face clusters.
Hugging Face had to fall back on a self-hosted Chinese openweight model, specifically GLM 5.2, too just to investigate their own breach.

We are not writing this because someone else had a bad week. Any team running agents at scale can be breached, and a security team that reconstructs what happened and says so publicly has done the industry a service. We have no interest in scoring a point off it.

The useful part is the shape of the problem, and it is a shape every one of us is exposed to. When something goes wrong, the question is not which model were you using. The question is can you reconstruct what happened — and the answer depends entirely on what you were writing down before you needed it.

That fear is the demand. So here is our own house, opened.

Our receipts

Every figure below was measured on our own box by the seat that wrote this page, and each carries the minute it was taken. The Brain grows on a 45-minute clock, so the counts move. We do not copy numbers out of last week's document.

resident AI employees, live6 of 6
source packets on disk4,979
documents in the retrieval index7,013
vectors · edges · documents with no vector10,132 · 1,867 · 0
units · scripture · canon · cards · catalogues4,976 · 980 · 877 · 178 · 2
sources with no resolvable raw file121 of 4,976
the whole index, one file254,115,840 bytes
billed today · list price of the same work$0.00 · $33.95
all-in running cost · ceiling$1.46 · $2.00 a day

seats, packets and index size measured 2026-09-03T16:35:21Z · index counts from the build at 16:10:19Z · meter written 16:43:25Z · cost banked 2026-09-01

Every source is kept as a packet: the verbatim raw text, where it came from, what it fed, and every pass ever made over it. Not a summary of the source. The source.

And 121 of them do not have that. Those units name a source file the resolver cannot find on disk, so the index counts them as raw-missing. That number sits in the same table as the flattering ones, because a receipt you only show when it helps you is not a receipt.

The gate that says no

This is the part we would want to see from a vendor, so it is the part we lead with.

A card in our Brain may not exist without naming the source it came from, and a quote in a card must be findable in that source's raw text. If it is not, the card is refused — and the refusal is written into the source's own file, where it stays.

sources whose proposed card was refused — the quotes were not in the source37
refusal entries in the recombination log, one per source37
of those still carrying the refusal as their standing judgment35
of the 37 that fired today33
of the 37 that are rehearsals, by registry membership0

window 2026-09-02T15:10:43Z to 2026-09-03T14:36:11Z · counted 2026-09-03T16:35:21Z · the most recent refusal fired 119 minutes before that count

Separately, the ledger that records refused card filings holds 42 rows: 22 production refusals — 21 cards naming a “source” that was not a real packet, and 1 filed without the name of the seat filing it — and 20 rows that are ours on purpose.

Those 20 are rehearsals. We feed the gate fabricated cards — invented quotes, sources that do not exist — to prove it still says no; 13 of them are invented quotations. They sit in the same ledger as the real refusals, each labelled REHEARSAL (synthetic), and every public count on this page excludes them by that label, never by guessing from the text.

A proof that a gate can refuse is worthless without a proof that the gate is still running — and a test you forgot to label becomes a lie in your own metrics.

The number that moved while we were writing this

Our drafts of this piece carried a retrieval result from the night before: ten of ten plain-English questions answered under one second, right card first on 9 of 10 questions. Re-running it at write time is a rule here, so we re-ran it. It did not reproduce.

Every run we took this morning is in this table. 25 runs of the same ten questions, 250 question-answers, none discarded:

right card first180 of 250
answered under one second41 of 250
fastest · median · slowest310 · 1,486 · 3,206 ms
runs clearing our own bar1 of 25
answers where the vector leg hit its 800 ms budget200 of 250

measured 2026-09-03T15:55:31Z to 16:31:14Z · the bar is 8 of 10 right and every query under one second · our own gate's verdict on 24 of the 25: fail

Why it moved, since a number without a mechanism is just a mood. Retrieval fuses full-text search with vector search, and the vector leg is given a hard 800 ms budget. If the local embedding model does not answer inside it, the vectors are dropped and the question is answered on full text alone — and the result says so, in a field built for exactly that:

"vec_error": "embed timeout — vectors skipped
             (an index build is holding Ollama); FTS-only answer"

The budget exists so a question is never answered late. The cost is that when the box is busy the answer arrives on time and less precisely, and this morning it was reached on 200 of the 250 answers.

Here is the part that did not go the way we expected. Eleven of the 25 runs recorded the box's one-minute load average as they launched: 0.97 to 5.51 on 4 cores. The three quietest of them — load 0.97, 1.06 and 1.12 — still put only 14 of their 30 answers under a second, even though a single search on that same quiet box at 16:40:29Z came back in 927.3 ms with its vectors intact. Load explains the worst of the spread and not all of it. We have not finished diagnosing it, and we would rather say that than publish the tidy version.

One more thing worth admitting. The organ that runs this proof overwrites its own receipt file on every run, so last night's passing result no longer exists as an artifact — what survives is a transcription we made of its values before re-running: 9 of 10 right, 10 of 10 under a second, slowest 894.7 ms. A receipt your own tooling overwrites is a receipt with a shelf life, and we found that out by needing it.

So the honest claim is not a single number, and we will not pick the flattering one: on this box, this morning, the median answer took 1,486 ms and 1 run in 25 cleared our bar. If a vendor shows you a retrieval figure without the load it was taken under and the runs that failed, you have learned something about the vendor.

Why we publish the misses

Because a number that can only flatter us is not a measurement, it is an advertisement — and we would rather be trusted in five years than impressive this week. We build as people who expect to be asked what we did with what we were given. You do not have to share that conviction to use anything on this page; the method works for anyone, which is rather the point.

The question we would ask you first

Not which model are you using.

Show me what your agents did yesterday, and what each claim was based on.

If you cannot list the actions, you do not have a system. You have a story about a system — and stories are fine until the week you need to reconstruct one.

That is the first thing our $995 assessment — five hours, or your money back — looks at. The Grill before it is free, and it is public evidence about your business, not ours.

If we have something wrong on this page, tell us and we will publish the correction.

Where every number here came from

Where every number came from

6 of 6 seats · 4,979 packets · 254,115,840 bytesmeasured 2026-09-03T16:35:21Z on the box
7,013 documents · 10,132 vectors · 1,867 edges · 0 without a vector · 121 raw-missingfrom the index build at 2026-09-03T16:10:19Z · the index reports its own stats
$0.00 billed · $33.95 listedpool ledger written 2026-09-03T16:43:25Z
$1.46 a day · $2.00 ceilingaudited and banked 2026-09-01, read not re-derived
37 sources · 37 log entries · 35 standing · 33 today · 0 rehearsalscounted across the source packets 2026-09-03T16:35:21Z; rehearsals excluded by registry membership
42 ledger rows · 22 production · 21 + 1 · 20 rehearsal · 13 invented quotesthe card-filing refusal ledger, counted 2026-09-03T16:35:21Z
180 of 250 · 41 of 250 · 310 / 1,486 / 3,206 ms · 1 of 2525 proof runs, 2026-09-03T15:55:31Z to 16:31:14Z, none discarded
800 ms budget · 200 of 250 reached it · load 0.97 to 5.51 on 4 coresthe budget is a literal in our search code; the counts and loads are from the 25 saved run receipts
927.3 ms on the quiet boxone search at 2026-09-03T16:40:29Z, load 0.75, vectors intact
9 of 10 · 894.7 msour transcription of the 2026-09-02T23:46:42Z run, kept because the organ overwrote the original
$995 · five hours or refund · free Grillour own frozen offer ladder, 2026-09-03

Market signal: reported, mined not adjudicated. Two independent public accounts, collided by the Brain on 2026-09-02T19:14:36Z and kept as units t5_009 and al_qPMhduk1qUs_001, each kept with its raw text. Every one of the three quotations above was re-checked against the raw text of the source that carries it before this page was rendered. No individual is named in this piece, and no claim is made about any company's spend, security posture or competence that its own public account did not make.