melchordgenscoreboard
← all reports

The milestone: any sample in, fire 2 of 3

The 17-test era is closed: the fire rate reads 8/17 (exact 95% interval on that rate: [0.230, 0.722], certified by the verification seat before printing; point estimate 0.47; the three takes per sample are not fully independent, so intervals read wider in truth). The two-in-three target is not met. The external review's "earned" benchmark (the interval's lower bound clearing 0.5) is not met. The measured trend: the era's second half keeps out-firing its first (+33 points at the last full measurement). A single passing sample still licenses very little on its own: the 95% interval on a 3-of-3 pass runs [0.29, 1.0], on a 2-of-3 pass [0.09, 0.99]. One pass carries the rater's own qualifier (thin, by his note, mood self-named as a covariate); one fail carries a reference score (the ground's own transcription rated 3, so the material set the ceiling).

The producer's question (July 9): how long until any sample produces a 4+ generation at least 2 times in 3? The measurement, under the bar he locked on July 10 (D203): sealed fresh samples, one exposure each, three regenerations per exposure; a sample passes at 4+ on 2 of 3, and pass rate and ground coverage are tracked as separate units. The coverage unit so far (8 fair tests): two clean grounds, a drone, and a near-repitch two-section ground (all passed, the last on the detector's closest borderline call, its adaptation praised at the ear); a dense in-key wash, a repitched chord hit, a moving progression with adjacent ground notes, and a wall-dense ground (all failed, each diagnosed; the sting class fixed and silent since, the wall ruled library-hard by his own reference score on the source). The repitch detector has correctly stayed silent on five straight non-repitch grounds. Forecasts on record: 12-18 batches (data seat), 15-25 (team lead); misses widen the interval. Full protocol on the methodology page.

One cell per exposure. Red = miss with its mechanism; the take scores are inside each cell. Passes carry 95% intervals of [0.29, 1.0] (3-of-3) and [0.09, 0.99] (2-of-3); the count is a footnote until 10+ fair tests. The three early misses converged to one structural law (L23: the sample's character drives the harmonic rhythm); the two later fails each carry a measured mechanism (a dense in-key wash; a repitch confirmed at +2 semitones).

Ratings history, b25 to the present (stage-relative eras, not one scale)

Best-rated clip per batch on the calibrated samples. Read this as a sequence of eras, not a trend: ratings are stage-relative (the bar rises as the engine improves), so a b25-era score and a current score are measured against different rulers, and the slope of this figure is not a measurement of anything. What the figure legitimately shows: which batches failed and why (red), and that the recipe reached repeatable 4s and 5s on known ground, which is why the milestone question became askable. Cross-time claims rest on within-batch blind A/Bs, not on this figure.

Scores are stage-relative: points in different eras are NOT commensurable and the line is drawn only to order the batches in time, not to show a trend. Red = failed batches (the rule-generation ceiling at b26, and the cases where a maximized metric lost to the ear); green = 5-rated wins. The dashed line marks the rising bar (D082).

Key detection accuracy

Auto-detecting a sample's key gates everything downstream (a wrong key rates low for the wrong reason). The early chromagram-era detector ran 36-60% on real loops. The current dual-method detector went 3 of 3 on b48's never-used keys and 10 of 10 on the held-out pool intake, with genuine ambiguities flagged to the rater instead of hidden. These later denominators are small (3 and 13, with 3 flagged ambiguities), so they are evidence of improvement, not a measured accuracy. One tension on the record, reconciled rather than overwritten: report 03 measured the older detector near 60% and named never-used keys as exactly where it missed; hours later the reworked detector went 3 of 3 including two never-used keys. Both are true; small samples bounce, and the 3-of-3 does not retire the 60% era's lesson until a larger denominator does.

Detection accuracy by era, denominators shown. The pool column includes three name-vs-sounding disagreements resolved in the sounding key's favor and recorded with both values.

Updated 2026-07-11, ~6:15 pm ET. This page fills in as measurements land; history is never rewritten (misses stay red). Exports: aggregates CSV · JSON feed.