Tag:evaluation
All the articles with the tag "evaluation".
One Prompt, Thirty-Two Calls, or Seven
Batching an LLM checklist coarsens the calls and forces the retrieval finer, in the same design. A seeded simulation of three call shapes, the pigeonhole that decides what a grouped call can even see, and the omission that used to come back as a severe finding.
Perfect Precision, Zero Recall, and No Curve to Trade Along
A document check cleared every forgery in a labelled set. The failure was the level of analysis, not the threshold: a seeded run shows the container detector reaching exactly two operating points, AUC 0.500, and shows what a genuine-only calibration can and cannot buy.
The Forecast Metric That Hides the Failure
A volume weighted error of 7.6 percent and a net bias of 0.2 percent can describe a forecast that loses to a seasonal naive on 77 of 116 slow moving series. The same weighting is applied twice more upstream.