Batching an LLM checklist coarsens the calls and forces the retrieval finer, in the same design. A seeded simulation of three call shapes, the pigeonhole that decides what a grouped call can even see, and the omission that used to come back as a severe finding.
A document check cleared every forgery in a labelled set. The failure was the level of analysis, not the threshold: a seeded run shows the container detector reaching exactly two operating points, AUC 0.500, and shows what a genuine-only calibration can and cannot buy.
A magic number survives because there is nothing to re-measure it against. The calibration set is usually already in a log you throw away, the same data sets the noise floor that makes any result readable, and the estimator you fix is rarely where the error was.
A volume weighted error of 7.6 percent and a net bias of 0.2 percent can describe a forecast that loses to a seasonal naive on 77 of 116 slow moving series. The same weighting is applied twice more upstream.
TabPFN-TS turns forecasting into a table-plus-features problem. A 24-hour edge case shows why geometry, checkpoint compatibility, and release evidence are separate gates.
LightGBM tuning can win, but only a deployment-faithful baseline and nested validation can show whether the last fraction of a point pays for its search.
Tabular foundation models are a real shift, but the durable bet is about information sets, task breadth, and synthetic priors - not easy benchmark wins.