News · triage

A daily agent that hands over a reading queue, and says what it could not judge

A lane can carry three live dossiers. Roughly a hundred and fifty plausible candidates arrive a week. The expensive decision is not what is true — it is what is worth a week, and this is the rule that decision follows.

Six free sources, six axes, four of them measurable by machine and two that need a person to read the thing. The machine fills what it can measure and reports what it cannot, rather than guessing a number to make the row look complete.

Last run 2026-08-05T12:04:20+00:00.

The queue

Ranked by the strict reading. Nothing here is open until a person reads it — the machine axes together describe a paper that is recent, on-topic, by named authors and carries a DOI, which is several hundred papers a week. The gap column is how many points are resting on judgment nobody has applied yet.

StrictMeasuredGapBandCandidate
58.990.6+31.7readThe Length of Martian Crater Rays and Their Relation to Lunar Cold Spotsarxiv:astro-ph.EP · 2026-08-03 · Trevor P. Erwin, Brandon C. Johnson, David Minton
55.685.6+29.9readSimultaneous Mars-orbit observations reveal Kelvin-Helmholtz instability-driven bulk atmospheric ion escapearxiv:astro-ph.EP · 2026-08-04 · Chi Zhang, Chuanfei Dong, Gangkai Poh
48.174.0+25.9logA study on the contribution of the interplanetary medium in radio occultation experimentsarxiv:astro-ph.EP · 2026-08-04 · Keshav Aggarwal, R. K. Choudhary, Abhirup Datta
47.873.5+25.7logNASA mission predicts solar explosion’s arrival at Earth with stunning precisionrss:science_news · 2026-08-04 · Rachel Berkowitz
46.271.0+24.9logSingle-Photon Counting CMOS Detectors for the Habitable Worlds Observatoryarxiv:astro-ph.IM · 2026-08-03 · Edwin Alexani, Justin P. Gallagher, Donald F. Figer
46.170.9+24.8logA new NASA Pioneer: the Globe Orbiting Soft X-ray Polarimeter (GOSoX)arxiv:astro-ph.IM · 2026-08-03 · Herman L Marshall, Sarah N T Heine, Alan Garner
46.171.0+24.8logHIP 61637 b: a TESS Brown Dwarf in a Near-circular Orbit around a Massive A-type Stararxiv:astro-ph.EP · 2026-08-04 · Nino Ephremidze, David W. Latham, Perry Berlind
45.670.1+24.5logDaily Briefing: How smallpox reached the Americasrss:nature · 2026-07-31 · Jacob Smith
45.670.1+24.5logESCAPE: a small explorer mission to study the stellar drivers of exoplanet evolutionarxiv:astro-ph.EP · 2026-08-01 · Allison Youngblood, Kevin France, Brian Fleming
45.269.5+24.3logCosmology in the Einstein Telescope era: comparing traditional and simulation-based methods for population inferencearxiv:astro-ph.IM · 2026-08-04 · Giovanni Antinozzi, Guillermo Franco Abellán, Davide Sciotti
45.269.6+24.4logThe Space Coronagraph Optical Bench (SCoOB): 11. Modeling and correction of chromatic aberrationsarxiv:astro-ph.IM · 2026-08-03 · Kyle Van Gorkom, Ramya M. Anche, Saraswathi Kalyani Subramanian
45.269.5+24.3logTomographer: End-to-end Redshift Distribution Estimation for Source Catalogs and Intensity Mapsarxiv:astro-ph.IM · 2026-08-04 · Yi-Kuan Chiang, Yu Voon Ng, Yu-Ren Lin
45.169.3+24.3logClimate benefit and ecological cost trade-offs for ocean iron fertilizationrss:nature · 2026-07-29 · Jun Yu
45.169.4+24.3logData reduction pipeline for the SuMAC millimeter-wave spectrometer at the LMTarxiv:astro-ph.IM · 2026-08-04 · A. M. Lapuente, P. S. Barry, M. Becerril-Tapia
45.169.4+24.3logFormation of nitriles and isonitriles by the heavy-ion irradiation of propionitrile in N2-rich astrophysical icesarxiv:astro-ph.EP · 2026-08-04 · Ana L. F. de Barros, David Dubois, Sreeja Raghunandanan

Denominator

A shortlist without a denominator is a recommendation. With one, it is a measurement.

BandMeansCountShare of the run
openstart a dossier00.0%
readworth twenty minutes of a person's time21.3%
watchreal, not this week00.0%
logrecorded, not pursued15398.7%
SourceCandidates
arxiv:astro-ph.EP25
arxiv:astro-ph.IM25
arxiv:physics.hist-ph6
federal_register14
rss:nature75
rss:science_news10

General news aggregators are deliberately excluded. A news feed delivers coverage-of-coverage, which is the material this rubric exists to reject — ingesting it would be paying to generate work for the filter.

The axes

AxisWeightKindRule
retrievable20machinea DOI, arXiv id or federal document number that a reader can pull themselves
lane_fit20machineoverlap with the subjects this show has actually covered, measured from the corpus
freshness15machinedecays over 60 days; a story nobody has touched in a quarter is not news
named_owner10machinea named author or agency who owns the claim and can be written to
decisive_test20humanis there an observable that settles it, runnable with the tools at hand
contradiction15humanvisible tension: claim against its own data, source against source, market against story
lane_fit is measured, not asserted. A hand-written keyword list would encode what someone imagines the lane to be. This one is derived at runtime from the scored claim corpus, weighting each topic by how central it is — 1,012 terms on this build. The vocabulary updates itself when new episodes are scored, so the lane never has to remember to tell the triage what it now covers.

Calibration — the test that can fail

The three dossiers already published were chosen by hand, months before this rubric existed. If the rubric scores them below its own open line, the rubric is wrong: the run exits non-zero and refuses to rank anything.

Current state: PASS

DossierStrictMeasuredBandWhy it was worth a week
trump-disclosure87.387.3openThree falsification observables with dated market settlements, and a priced gap between the corpus at 0.11 and Kalshi at 0.0605.
moon-ufo85.485.4openA pixel/sky/Moon null on the raw SER settles it, and the contradiction is internal — the label is refuted by numbers the authors published.
amazon-frontier81.481.4openA timestamped prediction scored against a later measurement; the tension is between detections in flown strips and a basin-wide extrapolation.
One result is worth reading closely. moon-ufo scores 0.28 on lane_fit — the lowest of the three, barely matching the show's measured vocabulary — and is carried by its two human axes. A purely mechanical triage would have filed it under log. That is the clearest available argument for why the human axes hold 35 of the 100 points.

This was not the first design. The first version let a candidate reach open on machine axes alone, and on the first live run three exoplanet papers were ranked as dossiers on the strength of having recent dates and DOIs. The band rule was changed rather than the weights, because the weights were not the thing that was wrong.

Two parser bugs, and the threshold they quietly invalidated

Recorded here rather than fixed silently, because the interesting part is not the bugs — it is that the output looked healthy for as long as they lasted.

The Nature feed supplies 75 of the 155 daily candidates. Every one of them was arriving with an empty date and an empty identifier. Not because Nature withholds them: the feed states both, as dc:date and as dc:identifier. The date parser tried three formats and Nature's plain 2026-08-05 was not among them, and the identifier was being scraped by regex from the link and description, where the DOI does not appear — while sitting in plain sight one element away.

Under the strict reading an unmeasured axis costs its full weight, so each of those items was docked 35 points of 100: fifteen for freshness, twenty for retrievability. Half the daily pull was being penalised for a missing format string. Nothing in the output said so, because a genuinely weak candidate and an unparsed one produce the same low number.

Fixing it changed the shortlist by zero rows — and that is the finding. Nature's best candidate rose 34 points, from far down the list to rank 40, and still did not make the cut, because it is about colorectal cancer and the lane is not. The ranking had been right for the wrong reason. What the bug had actually corrupted was the threshold: the watch line of 35 had been set against the depressed distribution, where it admitted 40% of candidates and looked selective. Against the repaired distribution the same line admitted 141 of 155.

So the line was re-derived from something that cannot drift with a parser: the four machine axes can award at most 65 points, and a candidate must now earn three-quarters of them — 49 — before it costs a person twenty minutes. That admits 2 of 155 today, or about 14 a week against a lane that can genuinely read fifteen; the agreement is offered as corroboration and not as the reason, because choosing a threshold for the count it produces is how a rubric ends up only ever agreeing with the person who wrote it. The reason is that an absolute line can report a weak day. A percentile admits the same share of whatever arrives and therefore cannot.

The federal-register lane was not affected and was not adjusted. Its documents score zero on freshness because the newest is dated 2024-12-11 — the UAP apparatus has filed nothing in twenty months. That is a measurement, not a bug, and it is the single most interesting thing this pipeline currently reports.

A third bug, and what the corrected axis then admitted

The parser fixes above repaired what the sources delivered. This one was in the scoring itself, and it had been inflating precisely the candidates the lane least wanted.

lane_fit tested each of the 1,012 vocabulary terms with a plain substring match. Nothing in the output showed what that meant in practice, so it was worth printing the matches and looking: sting was matching inside existing, pear inside appearing, lore inside explorer, craft inside spacecraft, and occult inside radio occultation — a standard astronomy term. A paper on interplanetary radio occultation was scoring as on-lane partly because the word contains the word occult.

Each false hit paid the same floor weight a genuine one-off subject earns, and the axis summed every match rather than the strongest ones — so eight accidents maxed an axis worth 20 points. Its own comment had always said "three central subjects present is a strong match", but the code divided by a literal 3.0 while only one term in the corpus carries full weight. No candidate could reach the ceiling through central subjects; it could only get there by accumulating generic words. The axis was measuring abstract length.

The threshold did not have to move, and that is the point of anchoring it to a principle. Word-boundary matching, the three strongest hits, and a ceiling read from the corpus instead of a literal 3.0 — three changes to the measurement, none to the line. Re-scored across the same cached pull, the read band falls from 21 of 155 to 2 of 155, and mean lane_fit from 0.240 to 0.173 (00_control_room/measure_lane_fit_change.py runs both scorers over the cache and prints this; the figures are that script's output on the 2026-08-05 pull). Under a percentile-shaped line all three corrections would have been absorbed by re-fitting the threshold, and the queue would have gone on looking exactly as healthy as before.

What the corrected axis then admits is worth stating plainly rather than tidying away. The three dossiers this lane actually chose to publish average 0.44 on lane_fit. The three candidates the same axis ranks highest today average 0.69, and 2 of them score above every published dossier. They are ordinary planetary science — crater ray lengths, an instability in the Martian ionosphere.

So the axis does one of its two jobs. It screens: 153 of 155 candidates fall below the best score any published dossier achieved, and the material it excludes really is off-lane. It does not rank: within what survives, being about Mars outscores having a decisive test, and the lane's own record says the second is what makes a story. An axis that works as a gate is being weighted as a score, and 20 of 100 points is the size of that mistake.

Not corrected here, on purpose. Re-weighting an axis against three published examples is fitting to three examples, which is the failure this whole page argues against. The test that settles it is dull and takes months: score each dossier as it is opened, and check whether lane_fit separates the opened set from the passed-over one at n large enough to mean anything. Until then it is recorded as a known defect with a stated size, which is worth more than a confident re-weighting.

Where this fits

Triage decides what to look at. The claim rubric decides what a claim is worth once you have looked. The dossiers are what comes out the far end, each with a call and a set of options. Full rule: RUBRIC.md, version 1.0.