Hundredfold Systems · field note · 2026-09-03

We already run the stack he described. Here is our meter.

A public interview this window walked through a multi-agent crew on a live knowledge graph — specialised agents behind one ingress, a memory fed from the channels a working person actually lives in. That is our category. We run it in production, and we can open the meter while you watch.

the meter · read from the box 16:35:21Z

$44

a month — every metered dollar the Brain spends. $1.46 a day averaged over 44 observed days, audited and banked 2026-09-01 — and it rides on a subscription already paid for other work, which we would rather say out loud than let you assume otherwise.

What that number does and does not include

The $44 is the marginal ledger, not the cost of this stack from zero. The ingest lane runs on a Claude Max subscription that was already being paid for other work, so its marginal cost is $0 and growing the corpus does not move the daily number. If you are starting from nothing, add the price of a subscription to our $44 figure.

We print that sentence because a cost figure that quietly excludes its own foundation is exactly the move we decline to make about anybody else.

model calls today484
billed$0.00
list price of that same work$33.95
standing daily ceiling$2.00

read from the box · pool ledger written 2026-09-03T16:43:25Z

The gap between $0.00 and $33.95 is a subscription, not a discount we invented.

What he described

market signal · mined, not adjudicated · a public interview ingested into our Brain as unit t4_048 on 2026-08-13 and carded by our Librarian on 2026-09-02

He runs specialised agents — Watson, Sherlock, Harry — behind a single ingress, over a persistent knowledge graph, with a private network between several machines. On routing, in his own words:

Watson then triages this and says, "Oh, this is a job for Sherlock." It's an image generation job.
Watson's also running a little relay service that is then sending um a a simple request to Sherlock over Tail Scale, alerting Sherlock to be able to do it. So I don't have like multiple web hooks looking at this.

Those quotations are ugly on purpose. The transcript is machine-generated and we do not tidy a quote. The stumbles are his, "Tail Scale" is the machine's spelling of Tailscale, and every quotation on this page is reproduced character for character, with nothing removed but the caption timecodes the transcript interleaves — which is why none of them carry quotation marks the source did not have. A quote you have cleaned up is a quote you have edited.

On cost, three fragments from one exchange. The transcript carries no speaker labels, and a short interviewer interjection sits between the first fragment and the second — so we are giving you the fragments rather than a sentence we stitched together:

at this moment I'm running all of this on an anthropic uh max plan
I'm not paying overages
I just pay for the electricity

And the contrast he draws himself:

if you hook it up to, you know, Opus, for example, via API, this is going to cost you like three probably 300 bucks a day

He never states what he actually pays per month. So there is no head-to-head on spend on this page and we will not manufacture one. What he describes is the same shape of economics we run on: a subscription, plus a local model doing the heavy repetitive work. The difference is not that we found a cheaper trick. The difference is that we publish the meter.

What we run, measured while this was being written

resident AI employees running6 of 6
source packets on disk, each with its verbatim raw text4,979
documents in the retrieval index7,013
units · scripture · canon · cards · catalogues4,976 · 980 · 877 · 178 · 2
vectors · edges10,132 · 1,867
documents with no vector0
the whole index, one file254,115,840 bytes
embeddingslocal, no API

seats and packets measured 2026-09-03T16:35:21Z · index counts from the build at 16:10:19Z · the packet count rises every 45 minutes as the ingest lane runs

Two of those numbers disagree on purpose. There are 4,979 packets on disk and 4,976 units in the index, because 3 packets landed after the last build. We are showing you the lag rather than quietly printing the smaller number twice.

Every source is kept as a packet: the raw text, where it came from, what it fed, and every pass ever made over it. Ask what an agent did and the answer is a file, not a recollection.

The runs that did not pass

This is the section that makes the rest of the page worth reading.

We keep a standing retrieval proof: ten real plain-English questions, each with the card that should answer it. The gate demands 8 of 10 answered by the right card first and every answer under one second. Every run we took this morning is in this table — 25 runs, 250 question-answers, none discarded:

right card first, among answers180 of 250
answers under one second41 of 250
fastest · median · slowest310 · 1,486 · 3,206 ms
the standing gatepassed 1 of 25
vector leg at its 800 ms budget200 of 250

measured 2026-09-03T15:55:31Z – 16:31:14Z · 25 runs of the same proof, every one kept

Our own earlier drafts of this piece carried “10 of 10 under one second, 9 of 10 right, measured 2026-09-02”. Re-measured at write time that does not reproduce, so it is not what we published.

Why it moves, honestly: the vector half of the search is capped at 800 ms on purpose, so that a question is answered on time rather than answered late. When the box is busy that leg times out, the search falls back to full text only, and it comes back slower and slightly blunter. It also says so — the response carries a field for exactly this, and the sentence it fills in is fixed in our source:

"vec_error": "embed timeout — vectors skipped
             (an index build is holding Ollama); FTS-only answer"

Eleven of the 25 runs recorded the box's one-minute load average as they launched: 0.97 to 5.51 on 4 cores. And here is the part we expected to go the other way. The three quietest runs — load 0.97, 1.06 and 1.12 — still put only 14 of their 30 answers under a second. A single search on the same quiet box at 16:40:29Z returned in 927.3 ms with its vectors intact (283.2 ms full-text, 641.8 ms vector, no timeout). So the load explains the worst of the spread and not all of it, and we are not going to pretend we have finished diagnosing it.

A system that can only report its good runs is not instrumented, it is advertised.

Where he beats us

A comparison that only flatters its author is an advertisement. Two places he is ahead today.

Reach. His connector layer, in his words, is "over a thousand different connectors including hundreds of models, you know, the ability to read and write from Twitter". We have nothing close to that breadth. It is our nearest gap on this axis and we are not going to pretend otherwise.

Hardware tiering. He spreads agents across several machines on a private network, one of them a gaming PC in his garage — "It's got an RTX 4080" — which lets him push whole classes of work onto free local inference. We run on one box with 4 cores. His arrangement is more resilient than ours and tiers work by cost in a way ours cannot today.

Where we are ahead

Provenance, per claim. Every card in our Brain names the source it came from, and every quotation is verbatim or the card is refused. That gate has said no in production, which is the only evidence that a gate is real.

The audit trail is the product. Not a feature bolted on afterwards, and not a log nobody can read. It is the thing itself.

The meter opens. You have just read our cost, our index, our speed distribution and our failing runs on one page. We would rather hand you the instrument than the summary.

Why we publish the misses

Because a number that can only flatter us is not a measurement, it is an advertisement — and we would rather be trusted in five years than impressive this week. We build as people who expect to be asked what we did with what we were given. You do not have to share that conviction to use anything on this page; the method works for anyone, which is rather the point.

If any of this is useful to you

The next rung is a $995 assessment — five hours, or your money back. We look at the channels you already have and tell you whether a crew like this over your inbox, calendar, CRM and socials is actually the move, or whether it is an expensive way to do something simpler. The Grill before it is free: public evidence about your business, no charge, no obligation.

If we have something wrong on this page, tell us and we will publish the correction.

Where every number here came from

Where every number came from

6 of 6 seats runningmeasured 2026-09-03T16:35:21Z · systemctl is-active, one call per seat
4,979 source packetsmeasured 2026-09-03T16:35:21Z · a count of the packet files on disk
7,013 documents · 10,132 vectors · 1,867 edges · 0 without a vectorfrom the index build at 2026-09-03T16:10:19Z · the index reports its own stats
254,115,840 bytes, one filemeasured 2026-09-03T16:35:21Z · the size of the index on disk
$0.00 billed · $33.95 listed · 484 callspool ledger written 2026-09-03T16:43:25Z
$1.46 a day · $44 a month · $2.00 ceiling · 44 observed daysaudited and banked 2026-09-01, read not re-derived
180 of 250 · 41 of 250 · 310 / 1,486 / 3,206 ms · gate passed 1 of 2525 proof runs, 2026-09-03T15:55:31Z to 16:31:14Z, none discarded
800 ms vector budget · 200 of 250 answers reached itthe budget is a literal in our search code; the count is from the same 25 saved run receipts
load 0.97 to 5.51 on 4 coresthe 11 runs of the 25 that recorded a load average at launch
927.3 ms · 283.2 full-text · 641.8 vectorone search on the quiet box at 2026-09-03T16:40:29Z, load 0.75, vectors intact
$995 · five hours or refund · free Grillour own frozen offer ladder, 2026-09-03

On his side of the page, every quotation is verbatim from unit t4_048 and was re-checked against the raw transcript before this page was rendered. He never states a monthly total, so no monthly total of his appears anywhere above. Our internal tracker filed this source under a company name that appears exactly 1 time in the machine transcript, against 6 occurrences of a different name in the same file. We caught it on the way to print. Nothing above depends on the name, so we left it out rather than publish one we had checked only once. Check the names your pipeline extracted before you publish them; ours nearly went out wrong.