Cases
These are not screenshots or write-ups of work you have to take on trust. Four of them are real HTML, CSS and JavaScript lifted out of the repository and running on this site against anonymised fixtures — the filters filter, the scores are computed by the same function that computed them in production, and where a layout had to be rebuilt rather than lifted, the case says so on its own page. The two arena boards are different in origin: they were built here, from published scientific catalogues, and each one says so in its own dateline.
Five of them are operating systems: something arrives faster than a person can handle it, so a machine researches it, scores it against a written rubric, and hands a human the decision rather than making it. The sixth is different — it is the claim ledger from the American Alchemy work, and it is here because it is the same method pointed at evidence instead of applicants.
AGI House runs dense founder and engineer hackathons, and admitting people, briefing sponsors, and later answering did this community actually elevate companies is not a spreadsheet problem. The engine scores each applicant on six role vectors — founder, engineer, research, PM, investor, media — produces a dated event briefing, and tracks founder impact over time with estimation labels attached rather than implied.
Those four numbers are four different denominators and the case page keeps them apart on purpose. The 13,733 is a count of a private corpus, published as a count and nothing else — no per-person record from it appears anywhere on this site.
Affiliate applications arrived faster than anyone could research them. An n8n agent researches each applicant and scores it out of 100 against a fixed four-factor rubric — reach 40, fit 35, engagement 15, identity 10 — and then, deliberately, stops. Every dossier lands in a review queue with a decision enum for a human to act on. It is a gate that is faster than opening ten browser tabs, and it is not an auto-approve.
Fundraising research collapses when it is ad hoc: no shared model of who is one degree away from whom, no way to hand a question to an agent, and no durable artifact at the end. This turned research this investor into a loop — task from Slack, agent run, committed profile page — with a first/second/third-degree circle graph underneath it.
The last figure is the point. The completed profiles are identifiable dossiers on real people, so the profile layout on this site is a reconstruction on invented data — and the case page says reconstructed rather than letting the reader assume otherwise. The original data file also carried a Slack channel id and real names; both were stripped for the fixture.
The NASA/USGS Mars Global Cave Candidate Catalog is a scientific record, not something you can browse and compare. Mars Arena wraps cave and lava-tube candidates in a sortable leaderboard with an inspector, scored by a weighted sum that lives in the page JavaScript where you can read it — safety 0.4, location 0.3, and the rest published on the case page. The score is a real function, not a decorative bar.
Zero network calls is deliberate: the CDN dependencies the prototype shipped with were replaced by a local sheet, so the case runs identically in five years and offline. The AI chat panel is still the placeholder the original HTML declared, and is labelled as one.
Eighteen lunar pits from the LROC Pit Atlas, scored on four axes. The board carries a second scoring mode whose only job is to show how much of the ranking rests on measurements nobody has taken: one reading averages the axes that exist, the other charges a pit for the axes that do not. The distance between them gets its own column.
Switching the reading reorders all but the bottom two, and the direction is not random — both fully-measured pits rise, all six two-axis pits fall. The two that rise are the two with published Diviner thermal data, which is also what the literature treats as the type examples. That agreement was not tuned for; it falls out of refusing to score a missing measurement as a zero.
The four systems above score applicants. This one scores claims. Out of 265.9 hours of canonical American Alchemy material it takes every statement made about Mars, classes each as fact, inference or speculation, scores the strength of the evidence behind it, and links the primary document or the exact video second that settles the check.
Seven threads are adjudicated against outside sources, and the Brandenburg material is split in two on purpose — the isotope observation and the civilisation conclusion are not the same claim and do not deserve the same confidence. Documents that were named but never retrieved are listed as named-but-not-retrieved, never rendered as if they were in hand.
The denominator is 136 canonical episodes, not the 172 uploads on the channel. The difference is re-uploads and clips, and quoting the larger number would make the coverage look thinner than it is while sounding more impressive.
Every one of them is a scoring function someone can read, applied to a working set someone can count, producing an artifact a human then acts on. None of them decide anything by themselves, and that is not modesty about the technology — it is the design. A system that auto-approves is a system whose mistakes you find out about from the person they happened to.
The other shared property is that each case states what is real and what was rebuilt. Lifting an interface out of a repository and running it on invented data is honest; letting a reader believe the invented data is a customer list is not.
Where a case shows a confidence band, the band is derived rather than asserted: eight components are scored separately and summed, so a reader who disagrees can disagree with one line instead of with a verdict. The rubric, the live distribution across 801 claims, and the three checks that audit the corpus against its own rule are on the method page — including the day one of those checks failed.
Client and partner work with no publishable artifact is not listed here or on the overview. A credential I cannot hand you a document for is not a credential.