Report 05 · July 10, 2026
The day the engine earned its first fires, lost its ghost, and hit the milestone. Four morning misses became a law; the off-time complaint that survived seven hypotheses turned out to be one hardcoded 104 bpm stamp, found by his question; and the first measurement through the fully corrected pipeline came back 5, 5, 5 on a never-seen sample, the next fresh exposure fired again at 5/5/3, and the third missed one take short of the gate, handing over a fourth character axis (density) with its mechanism measured. Clean era: fire, fire, miss. The milestone needs a fresh three-streak, with 21 counted samples to run it on.
TL;DR
Exposure one is live on the rating app: crystalpeak, A minor, a key the engine has never generated in, three independent full-stack takes against an absolute bar (fire = 4 or better; the milestone datapoint = 2 of 3). The chain that delivered it ran its first live cycle with zero deviations, including a real contamination catch that shrank the pool from ten samples to six after the provenance screen.
The state carried across midnight, as a table:
| component | state | since |
|---|---|---|
| generation engine | clean lineage (connections + gold + blend adapters), weights byte-verified local | b47 tie, adopted Jul 9 |
| blend depth | STANDARD (1.00) default; HEAVY (1.30, mined) ships; LIGHT (0.85) internal yield tool | b53-b55 trilogy |
| take selector | context-adaptive band-fit | b49 tie |
| voicing | rule-builder | b50, decisive |
| chord initiation | adaptive grid (L22) | b51, decisive |
| plugin ports | 4 research pieces in C++, python-parity tests in the gate | T032-T033, Jul 9 |
| library | 8 style cells, ~24 keepers, internal | Jul 9 |
| GPU pod | terminated; all generation runs on the local 3080 | Jul 9 |
| pace forecast | 12-18 batches (data seat) / 15-25 (team lead), ungraded | Jul 9 |
Yesterday's full story, 23 timestamped entries, is report 04.
The first measured datapoint is in, and it is a miss: 2, 2, 3, zero fires, with the failure mechanism named. His ear named it twice, unprompted: the generated chords fighting chords already in the sample. Chord-on-chord collision. crystalpeak carries its own moving harmony (F, E-minor, A-minor, G), and the engine free-generated its own chord bed in the detected key without following the sample's bar-by-bar changes. The measurement corroborated his ear exactly: 30 to 32 out-of-harmony pitch collisions per take.
Two findings make the miss valuable. First, the blind spot was structural: both trusty calibration samples are sparse loops, so a chord-bearing source could never expose this failure in fifty batches of tuning, and the held-out protocol found it on sample one, which is precisely what it exists for. Second, the generalization split, in his rating notes, the way the generation moves with the sample drew real praise. The timing and groove layer (adaptive initiation, the bands, the interplay mechanics) generalizes to never-seen ground; the harmony layer fails specifically on harmonically dense sources. The hardest-won layer works; the gap is localized and named.
And the protocol held under its first disappointment: the forecast interval widens per the pre-written rule, crystalpeak burns to calibration, and no same-day fix ships. Fix development on burned material is legal and immediate: a chord-following layer (variance-gated so sparse loops keep the free-gen path that earned the fives, recovering the sample's literal root movement, voicing from the detected set) already cuts collisions from 30-32 to 6 on the burned sample. Validation happens on the next fresh exposure only; the next lever came from the collision measurement rather than a hunch.
Exposure two (MELODYtool, 83 bpm, G-minor sounding) rated 3, 3, 3: zero fires, second miss. The difference from exposure one is the point. Before his ears touched it, the routing decision, the sample reading (a G pedal), and the risk were preregistered in the committed metadata, stated both ways so the outcome could not be spun.
Two things resolved. First, the chord-following fix held on its first blind test. His note on the takes, paraphrased: super stagnant and a bit annoying, though not incorrect, which he counted in its favor. The chord-on-chord wrongness that sank exposure one is gone on chord-bearing fresh ground; on the burned calibration clips the wired path measures 4-6 collisions per take against 30-32 for free generation, an ~80% reduction, and now his ear corroborates it blind. Second, the miss matched the preregistered risk exactly: following a one-chord pedal is correct but static, and static is what he heard.
His verdict also named the next lever, asking for an occasional different chord arriving later: harmonic development across the bars instead of maximal continuity. The detected bar-sets on this sample contain secondary root candidates (D, B-flat, C), so the material for development exists in the sample itself. Next actions, per protocol: MELODYtool burns (the two burned samples now form a two-ground calibration set), the root-development fix (T038) builds on burned material only, and validation happens on exposure three. One engineering note recorded with the wiring: the sequence avoid-list (L21) is deliberately skipped when roots are followed from the sample, because it governs the engine's own generative choices, not fidelity to the producer's source material.
Update, 9:45 am: T038 is built and exposure three is packaged. The development fix carries his request written into its docstring and a correctness invariant: every root still comes from the sample's own detected per-bar sets, and a bar develops only after the same root has repeated twice. On the two burned calibration samples: crystalpeak is unchanged (the cap never binds on moving harmony, so no regression), and the MELODYtool pedal develops from one distinct root to three, in the back half, which is where he asked for it. The routing gate then produced a correct surprise: the next sealed sample (MissYou) measured as a static loop (variance 0.143) despite its chord-bearing vibe, so it routes to free generation. The preregistered consequence: exposure three tests the free-gen path's generalization on fresh static ground, and develop mode waits for a chord-bearing exposure. Held for the delivery precheck.
Exposure three (MissYou, 100 bpm, C-sharp-major sounding, routed to free generation by its own variance measurement) rated 2, 1, 3: zero fires, third miss. His notes carried the diagnosis: "too changey for this specific sample... since the sample sustains same chord you can do a dif chord like a measure or two later," alongside "chords feel pretty good though" and "really close with the chord voicings, just timing not great."
Put the three verdicts side by side and they stop being three problems. On the moving sample the engine ignored the changes (wrong). On the pedal it froze (boring). On the drone it changed every bar (too changey). One root cause in three costumes: the sample's character must drive the harmonic rhythm, and the engine had no dial for that. That is now L23, the most heavily evidenced law in the taste bible: measured on three independent grounds, with his design instruction embedded verbatim. The generalization split also sharpened: across all three fresh samples his words approve the content (chords good, voicings close, movement great) while the structure, when harmony changes relative to this sample, missed every time.
The response (T039) turns the variance number the pipeline already computes from a binary gate into a continuous chord-change-rate dial: a drone gets one or two changes per eight bars, arriving late, which is his instruction taken literally; the trusty band keeps classic two-bar spans; moving harmony keeps per-bar following. Calibration required all three burned samples to route correctly simultaneously, with one dial and zero special-casing, and they do. The cadence call was his to make (pool consumption is owner territory) and he made it: his deferral to the recommended option Exposure four (Gun Barrel, C-sharp-minor sounding) is generating now as the dial's first blind test, with the character reading and full change schedule preregistered before his ears. One administrative flag rides along: three clean samples remain after this one, past the pool's refill trigger, so the refill ask (5-10 fresh loops in the same folder) is on his plate.
Update, 10:45 am: exposure four is packaged and holding at the precheck. The gate read Gun Barrel as a drone (measured, not assumed from its name or key confidence), and the preregistered change schedule is two chords over the window with the change arriving as designed. The calibration evidence also became a committed file rather than a comms message, per the git-versus-disk rule. His fourth listen follows the precheck.
The verdict: 3, 2, 2, zero fires. The milestone scoreboard reads 0 of 4. And inside the miss sits the morning's biggest result. His note on the best take split cleanly in two: the chords read as genuinely good; the timing read as odd, unmusical, close to random. Read those halves separately, because they grade two different levers.
The chords landed. The character dial's schedule for a drone (hold, then one change arriving late) drew approval from the ear that called exposure one wrong, exposure two boring, and exposure three too changey. None of those three failure modes recurred. L23's harmonic half, the lever three misses built, held on its first blind test, on never-calibrated ground (one ground; the claim stays this size until more).
The timing did not. His other note placed the strikes on off-beats that never connect to the sample, worse in the later bars. That complaint is not new: exposure three's weakest take drew the same random-timing read on the other drone. Two independent static grounds, the same words. The miss has moved to a new and narrower address.
One precheck detail deserves its line: QA flagged the source as too quiet before delivery, and the fix was a pure gain multiply (×7.1, to the loudness of the first three exposures) applied at delivery only, with the sealed source untouched and the change noted in the exposure's metadata. The b48 postmortem class (the inaudible-sample failure) was defended before it could contaminate a datapoint, not diagnosed after.
The diagnosis behind the new lever is a scope lesson, not a broken measurement. The chord-initiation grid that won b51 (the emphatic first-listen yes, law L22) places chord strikes in the melody's and sample's rhythmic gaps. Its calibration bands (gap-fraction 0.28 to 0.51) were measured from his liked tracks, which are full mixes with drums. Against crystalpeak's produced groove the gap-answers bounced and his note praised the movement. Against a drone there is no groove to answer, so the same off-grid strikes read as random. The measurement was right; its scope was narrower than anyone knew until two drones said so.
L24, minted on two grounds: chord timing is character-contextual. Grooving samples get gap-adaptive strikes; floating samples need metric anchoring, strikes on beats the ear can find. This extends the character dial from the harmonic axis (how often chords change, L23) to the rhythmic axis (where the changes land). The burned quartet now holds two drones, which makes it the timing lever's calibration pair; the lever develops during the interlude and tests blind at exposure five.
Four datapoints in, the ledger of what travels to never-seen ground and what does not:
The travels ledger: each engine layer, whether it has held on fresh ground, and the evidence. Every row cites his words on a sealed, never-calibrated sample. The n in each cell is the number of fresh grounds the claim rests on; a capitalized verb without its denominator would be a claim wearing the costume of a conclusion.
A design entry from the owner himself: the classic functional-harmony chart (tonic, predominant, dominant) with his framing attached: functional harmony as pulls with weights, not boxes in a fixed order, worth testing eventually with the weighting itself on trial. Banked to the repo as a design note with a concrete integration map: first ranking the drone change-chord, then a development bonus in the root-follow layer, then a soft term in free generation.
The measured caveat that keeps it a prior rather than a rule: his corpus is oscillation-dominated, and his rated winners include non-functional moves he explicitly told the team to keep. So gravity enters as a soft prior, the corpus wins conflicts, and the weight ships only through a blind A/B. Theory proposes; his ear disposes.
The interlude's first hour went at L24's open question: what actually separates ground where gap-answer chord strikes bounce (crystalpeak, where his notes praised the movement) from ground where they read as random (both drones)? Four candidate features went in. All four came out dead, and the deaths are the data.
Hypothesis one, sample groove strength, refuted by its own table. Onset rate, onset strength, and pulse clarity were measured on all six known grounds and none of them separates the ear's verdicts:
Hypothesis two, the rhythmic density of the result, killed by a free experiment. The idea from his unmusical-metric complaint: dense melodies carry the meter themselves, making off-beat chord answers read as syncopation, while sparse melodies leave the chords carrying rhythm alone. The test, b56, cost zero pool: trusty D-major ground, one deliberately sparse melody (8 notes, full screens) shared by both arms, and the only difference when the chords strike: adaptive gap-answers versus metric strong beats. Blind, QA-passed 2 of 2. His verdict: 4 and 4, no audible difference. Sparse texture alone does not reproduce the complaint.
Hypotheses three and four died by measurement before reaching an ear: MissYou's pads crest like crystalpeak's drums (attack crest 5.07 vs 5.16, no separation), and MissYou carries the highest percussive-energy ratio of the set (0.095 vs crystalpeak's 0.037) while feeling the most random. All four feature tables are committed so nobody re-treads them.
The ruling is the project's first risk-profile resolution: no feature gate, an expected-value decision. Metric strikes are non-inferior on rateable ground (the tie), while adaptive has helped once on fresh ground (the drummed sample) and hurt twice (both drones). So exposures now default to metric strikes (wired the same hour), trusty lever batches keep the adaptive feel b51's verdict earned (the 5 stands, no churn), and adaptive-on-fresh returns as an earned upgrade only when something real separates drummed from floating material, likely a learned percussion classifier, now that four hand-built features have failed. The self-check clause is preregistered: if the timing complaint recurs on metric at exposure five, the diagnosis was wrong at a deeper level and the report will say so.
Within minutes of the timing ruling, the owner corrected its framing, and the correction is ledger-grade: always relying on the same progression or the same harmonic movement defeats the point, because the goal is context-awareness. The reframe is committed: the metric-timing default is a floor, and it exists only because four context-detectors failed measurably. It is not the architecture, and it never graduates into one. Context-awareness stays the goal.
The correction generalized into a proposed engine principle, now with the team lead to shape: every uniform default must carry (a) a documented context-upgrade path, (b) the evidence that would activate it, and (c) an expiry review under the standing ten-batch drift check. The one-line version from the team's own note: a default without an upgrade path is context-blindness with paperwork.
Two threads opened directly from his correction. First, CLAP as a groove detector is untried and in scope: the audio-embedding backbone's style-level judgment was the territory its backtest left it (only its take-ranking died, at 0 of 6), and a cheap probe is queued for an interlude slot. If it separates the six known grounds where four hand-built features could not, the timing axis gets its context back early. Second, content variety on exposures: free generation's concentration habits (vamp families) hold the exposure incumbency by default rather than by fresh-ground evidence, so the progression-source axis (model versus retrieval versus follow) gets an exposure-side revisit once timing settles.
Update, just after noon: the CLAP probe ran, and it failed too. Zero-shot CLAP reads MissYou as the grooviest of the six grounds while the ear feels it as the most random. That is null-result entry five (groove features, result density, attack crest, percussive ratio, CLAP zero-shot), all tabled in public. The metric floor stands with an upgrade path that is, as of noon, empty of candidates, and the floor principle itself is now minted into the taste bible.
The day's meta-arc, in one line: the morning's drift-fear note and this correction are the same instinct, and both are now standing rails. The engine stays accountable to the goal of listening to the sample in front of it, not to its own habits.
The owner asked the question the protocol was quietly avoiding: why exclude samples the model trained on, when what matters is whether the engine generates well over them, complementary material rather than copies of those melodies. The answer that came back splits the concern cleanly, and it is now a proposed protocol extension awaiting his word.
Same-source samples can never count toward the milestone: training saw those tracks' derived content, so a fire there would measure memory, not skill, and the fixed star would silently move. But testing them is a ceiling measure with real diagnostic teeth, on a second scoreboard that is labeled as such and kept separate. The readout logic: fires at home plus misses on fresh ground = a pure generalization gap, and the forecast holds its shape. Misses at home too = the problem is deeper than generalization, and the forecast gets rethought. The rateable-never-counted tier QA built into the pool design was made for exactly this and has been unexercised until now.
The proposed first tier-2 exposure doubles its value: the the kart chord stem (key confidence 0.923, clearly chord-bearing), which would also give the follow and develop paths the blind test ground they still lack, at zero cost to the clean pool. Two scoreboards, both labeled, neither contaminating the other. It runs on his word.
Update, early afternoon: he said yes, and tier-2 exposure one is packaged, holding at the QA precheck. The preregistration carries the project's first borderline routing case, shipped exactly as the machine decided it: the kart stem's harmonic variance measures 0.573, which is 0.027 under the follow threshold of 0.60, so the engine routed it to free generation. The committed collision report predicts that decision is wrong: 14 to 16 semitone clashes per clip, the worst of any exposure yet (exposure one's failing takes had 4 to 6). Per the anti-retrofit rule, nobody retunes the threshold mid-exposure. The prediction on record before his ears: if he hears chord-fighting, the follow threshold drops below 0.573, a lesson bought on the home-field board at zero milestone cost; if he rates the takes well despite the collisions, the variance metric itself needs rework. Either verdict teaches something the clean pool would have charged for.
His verdict on the kart stem: 2, 3, 2, zero fires on the home-field board. The milestone is untouched by construction; this scoreboard was built to buy lessons, and this miss bought three.
First, the borderline prediction confirmed at the ear. His note flagged the chords as clashing and wrong, which is exactly what the committed collision report said would happen when the machine routed a rich moving chord stem to free generation. So the threshold learned from his ear, not from a retune: the follow line moved from 0.60 to 0.50, wired the same hour. the stem's 0.573 now routes to follow; the trusty ground's 0.429 stays free-gen, untouched.
Second, the self-check clause fired. The off-time complaint recurred on metric strikes, all three takes. The clause written into the b56 ruling said exactly what this means and the report says it plainly: the metric floor does not close the timing class, a seventh probe (beat-phase agreement) also failed (the trusty ground phases at 67% with zero complaints, so the tracker lies on sparse loops), and as of this verdict the timing failure mode was not understood. His stated worry got the state of the evidence, unspun: six failed detectors on the table in public, candidates recorded but not chased, one mechanical probe per candidate before any batch spends an ear.
Third, the ceiling read is confounded. This sample got predictably wrong chords, so a home-field miss establishes nothing about the ceiling yet. The clean re-run after the fixes is owed and on the board.
Six minutes after the "not understood" admission, the seventh probe came back separating cleanly. Pulse clarity at the working tempo splits all six known grounds: every ground whose timing drew complaints measures under 2.0, every clean ground measures 2.75 or higher, and the detector line sits at 2.4 with real margin on both sides.
The probe run also killed the tempo-mismatch framing on the way through (the engine's tempo estimates are right), and its values forced the reframe that matters more than the detector: the off-time complaint was never about where the strikes landed. Weak-pulse samples give the ear no shared clock, so no strike schedule, adaptive or metric, can lock against a grid the sample does not state. The complaint was unanswerable by any of the levers tested, which is why six hypotheses died.
The fix direction is a new class, not a tweak: grid-establishing generation on weak-pulse material, where the piano states the time itself, steady and meter-carrying, instead of answering a groove that is not there. Harmony still changes on the character dial per L23; clear-pulse samples keep the metric floor. And the ear test is already the natural next event: a tier-2 re-run on the burned stem Summit carrying both of the day's fixes (the 0.50 routing now sends it to follow; the timing now states the grid), on the home-field board, at zero milestone cost. It is generating now; it reaches the app on the team lead's shape and his one word.
Update, 1:48 pm: the survivor earned a caveat within the hour. After the intake cut bug was found and fixed (pane 14), the corrected cut flipped the stem's pulse-clarity from weak to clear (2.71). The detector separated the six grounds partly because it was reading the project's own cut misalignment. The grid-establishing class stays parked; the re-test plays standard strikes on the corrected cut.
Tier2-02, the kart-stem re-run with both of the morning's lessons aboard, came back 2 of 3 at fire on the home-field board: the measurement era's first fires anywhere. His notes called the chords and voicings the real thing, held back only by the timing: these would be 5s if they sat in time. The follow routing at 0.50 did what the preregistration predicted, collisions went from 18-20 per take to zero, and the roots are the sample's own B-minor family. The milestone scoreboard is untouched at 0 of 5 by construction; this is the ceiling board lighting up on the chord axis.
And the qualifier in his verdict solved the day's central problem. The grid-establishing takes play steady metronome quarters, and he still heard them as clearly off. A metronome cannot be musically off-time. So the fault could not be in what the engine plays, only in where the audio sits, and the alignment probe confirmed it on all three complained grounds: the intake cuts land 11 to 18 percent of a bar off each track's own grid (that stem by 248 milliseconds). The mechanism is a plain bug: the intake window slides from time zero, assuming the audio starts on beat one, so any intro or pickup offsets every cut, and every strike of every kind lands off the sample's true beat.
Seven musical hypotheses died hunting a cut bug. It is the b48 class one layer deeper: a delivery-adjacent fault masquerading as an engine failure, contaminating the timing component of every exposure verdict to date. The datapoints stand, nothing is retro-scored, and the diagnosis is reframed. His ear found it, twice over: the "metronome cannot be off-time" deduction only existed because he rated a metronome as off.
The fix shipped within the hour: intake v2 cuts from each track's own downbeat instead of assuming beat one at time zero, in both frozen-layer copies, versioned, with QA re-verifying. The full-pool probe confirms the mechanism and the ear agree: the three timing-complained grounds are the three worst cuts (11 to 18 percent of a bar off-grid), MissYou is the worst in the pool at 27 percent, and every clean-rated ground sits inside the 5 percent noise guard, where a shift that small reads as onset lag and the frozen cut is kept.
The probe also handed hypothesis seven its caveat, reported as plainly as its survival was: the corrected cut flips the stem's pulse-clarity reading from weak to clear (2.71), so the detector was partly measuring the project's own cut misalignment, not only the sample's pulse. Tier2-03, the same stem regenerated on the corrected cut, therefore plays standard fixed strikes, is packaged with its preregistration (rating-batches, tier2-03), and awaits the QA precheck. Both readings are pre-committed: timing clears and alignment was the wall, with exposure five running on the v2 pipeline after his refill; timing persists and the residue moves to app playback offset, a QA lane, explicitly not a ninth musical hypothesis.
The hour after the root cause was about closing the seam it slipped through, and it started with QA owning the gap out loud: the window-sync check gated cut duration but never downbeat phase, which is exactly where the bug lived. The phase check is built to close it, and QA requested an independent alignment number from the probe before anything reaches his ears again. Every seat has now self-caught something today.
The post-fix instrumentation reported its own limits along with its numbers. The chord-change phase concentration drops from 46.7 to 29.3 percent after the corrected cut, and the v1 signature it replaces is a half-bar offset, the measured shape of his off-time complaint. A second instrument (the energy comb) was tested and discredited. And the limit is stated plainly: no instrument certifies alignment below 5 percent of a bar on reverbed ground, so his ear stays the decider, exactly as the preregistration assigned it.
Hypothesis seven got its final status: re-measured on the corrected cuts, two of the five grounds flip from weak to clear, the weak-pulse class shrinks to MELODYtool alone, and the detector dies with credit: it separated the grounds partly because it was reading the cut bug, which made it the arrow that pointed at the real fault. Tier2-03 is prechecked green (QA's seventh consecutive) and holds for his word; a separate drift probe came back null, so if timing somehow persists on the corrected cut, the escalation goes to app playback first, not to another musical theory.
Tier2-03, published under a new standing delegation, now ledger law: rating-app publishes no longer wait on his word, while pool consumption stays his. It came back 2, no fires: off-time persisted on the corrected cut, plus a new first-chord dissonance. QA logged it as a confounded datapoint (the new cut changed the generation, tangling a voicing regression into the timing read). And his note carried the key: he asked whether the tempo had really been checked.
The tempo detection was right. The tempo stamp was not. The MIDI writer has hardcoded 104 bpm into every batch since its D-major birth, and the rating app clocks the piano from that stamp: the piano plays at the wrong speed on every non-104 ground, drifting further off all loop, snapping back at the seam. The correlation across every timing verdict of the era is total:
Entry 8 is corrected in place, not quietly replaced. The metronome that "cannot be musically off-time" was a 104 metronome playing against a 102 sample: it was off, his ear was right, and the cut misalignment it exposed is real but exonerated for timing (the pickup trap it found becomes intake v2.1, a harmony-hygiene item, after tier2-03’s bar-one dissonance traced to a C-sharp tail the earlier cut pulled into the first bar’s pitch set). Eight musical and mechanical theories, two of them shipped as fixes, and the fault was a constant in a writer function. The fix stamps the holder’s true tempo, in both frozen copies.
The cleanest experiment of the project: tier2-04 is tier2-02’s exact fire clips, byte-patched to 102.000 bpm, zero regeneration. Same notes, same voicings, same everything, only the clock corrected. His verdict: 4, 4, 4. Home-field 3 of 3 at fire, the first full sweep. His notes read the takes as on time and good when synced. Music he rated 2 an hour earlier went 4/4/4 with only the tempo stamp changed. Entry 9 is confirmed at the ear, and the timing era closes.
The consequences reach the whole scoreboard. Every timing complaint since exposure one reframes as clocked-wrong delivery; the harmonic verdicts were never touched, which is why L23 and the follow threshold survive intact. Exposure five, whenever his refill or his call spends the 135bpm loop, becomes the milestone’s first clean measurement: true clock, the follow stack, and a new axis he directed into batch design on the spot: vary the progressions, and the starting chords too. The variety axis: one take follows, the others take L21-clean alternative progressions and starting chords. The residual drift is bounded and understood (at most 18 milliseconds per pass against the ~102.10 track, resetting each seam), QA gets one instrumented pass on the live app as hygiene, not a gate.
The hour after the sweep turned his directives into pipeline. The variety axis shipped: on chord-bearing grounds, take one follows the sample purely while takes two and three play bar-consonant reharmonizations with distinct progressions and distinct starting chords, all L21-clean, with the exposure-one collision lesson as a hard constraint. Intake v2.1 shipped as a purity duel: a cut anchor must now earn any move by beating the frozen cut on bar-harmony purity, which restored the stem’s fire-rated cut and left every other ground untouched. And the fine-tempo pass closed the last decimal: batches now stamp the fine-scanned rate, killing the bounded 18-millisecond residual.
Update, 3:07 pm: the precheck became a program. QA’s delivery gauntlet is now one script with nine checks (seal, cut integrity and audibility, the full screen pass, tempo stamp against fine-scan, downbeat phase, preregistration-both-ways, sounding key, bpm, and the em-dash rule), emitting green or red plus a committed audit report per batch. Its first live run: exposure five, green, zero warnings: stamp 0.000 off the fine scan, downbeat phase at 3.8 percent. Publishes now ride script-green under the standing delegation, QA holds countersign and instant-rollback authority, and the milestone’s first clean measurement went to the canonical app.
An owner-requested comparison to close the arc: one clip from b25 (July 6, the first batch a model trained on his own music ever produced) against one clip from tier2-04 (July 10, the batch that just swept the home-field board). Same engine lineage, four days of verdicts apart.
| b25 · July 6 | tier2-04 · July 10 | |
|---|---|---|
| his rating | 3 (beat the old baseline’s 2) | 4, on all three takes: the first full sweep |
| what the model was | one melody adapter over a hand-assembled chord bed | the three-adapter engine (cohesion, his sound, loved movement) + likes-trained progressions |
| what it knew about the sample | key only | key, tempo, harmonic variance (routing at 0.50), character (chord rate per L23), its own chord roots followed |
| what stood between model and ears | nothing: raw takes, hand-picked | eight screens, collision pass, QA precheck, preregistration before his listen |
| timekeeping | the 104 stamp, silently wrong off-104 grounds | true clock, stamped from the sample itself (Entry 9) |
| training data | 19.6k of his melody clips, key-agnostic, pre-dedupe | the deduplicated corpus (11,461 real pairs) + 120 liked-track progressions |
The methodology gap is the story of the whole site compressed: b25 was a model with nothing but taste in its weights; tier2-04 is that taste plus four days of his verdicts turned into routing, screens, laws, and a corrected clock. The full engine map, adapter by adapter with the future layers marked, lives in report 04, the adapters entry.
The milestone’s first fire-datapoint is a perfect sweep on fresh ground. a 135 bpm loop, a sample the engine had never seen, through the first fully clean pipeline: true clock (fine-scanned 134.92), purity-guarded cut, the follow stack, grid-establishing strikes. His verdict: 5, 5, 5, capped by the note the whole effort aims at: he would use this. The thesis's first confirming datapoint: one never-seen sample, one listener, one sitting; the exact 95% interval on a 3-of-3 pass is [0.29, 1.0]. [Calibrated 2026-07-11: this line originally read "the thesis, landed", which one datapoint does not support.]
The preregistration scored clean on every branch: the in-time read at 135 bpm confirms the old 30-percent-error class is silent; grid-establishing is redeemed at 5s, its first fair test since the confound (the weak-pulse treatment now has its first supporting verdict); the texture-smear risk never materialized. His one flag, on clip two, diagnosed mechanically within minutes: the melody’s final note sustains across the last downbeat as an unresolved ninth over the closing chord. In key, passes every screen, purely an ending-resolution gap, and it becomes the next lever: final melody note resolves to a chord tone or releases before the last downbeat (a warn-class screen plus a selection preference, held at n = 1; the 5 stands).
The scoreboard, worded per the team lead: datapoints one through four, no fire, with their timing component reframed by Entry 9 and the datapoints standing; datapoint five, fire, 3 of 3. The clean-pipeline era is 1 for 1. the 135bpm loop burns; Elevator is the last clean sample; the refill ask stands.
Update, 3:30 pm. The fire's paperwork landed: the ending-resolution lever is cut as a task (warn-class screen plus selection preference, held at n = 1), and the refill is now formally the milestone's only bottleneck. One more catch chain rode the hour, and it shows the process working in all directions: the owner spotted a stale batch title on the rating app, the fix went out, and QA then caught the fix itself incomplete (two more stale strings), so a generic guard, zero hardcoded batch text plus a foreign-label scan, is now demanded before the checklist countersign. The verdict is unaffected, the audio was correct and the exports prove it, and the fix goes forward, not backward. For the record the pace deserves: the first scored point on his any-sample thesis landed about 26 hours after he asked how many batches it would take.
The first thing the owner did with a working engine was tell the team what it becomes. The vision, banked the same hour as a committed design note: self-distillation. Once the engine is trustable at the milestone bar, it farms its own paired training data headless, nights and days, generating over samples and keeping what survives the screens. On top of that dataset: a two-tower, audio-conditioned student, a model that hears the sample directly (both the audio and the MIDI of every pair, because the spectral and timbral context is part of what must be understood) instead of reading detected features.
Two things make the note engineering rather than a wish. First, he spotted the failure mode himself, unprompted: model collapse, the known degradation when models train on their own outputs. The design answer rides in the note: synthetic data stays the minority against a gold anchor of real pairs, audited in small batches. Second, the sequencing guard: milestone consolidation comes first, the endgame task stays parked, and the fixed star does not move for it. The gold anchor itself: his own project archive, where he has been playing over samples for years.
The feasibility question got numbers the same afternoon. A probe walked all 1,485 projects with zero parse errors hunting exactly what the endgame needs: sections where unmuted audio of real length overlaps his MIDI, him playing with a sample. It found them everywhere.
The sizing, deflators included, is the story. Two of the owner’s own catches, made mid-conversation, were load-bearing and are now quantified. His pitch-truth catch: 37 percent of pitched paired clips (22,179 of 60,657) ride a Pitch MIDI effect, plus 760 warp-transposed audio clips, so a third of the anchor would have carried wrong-key data without the correction. His MIDI-first rule: 53 percent of the raw clip count was Sampler and DrumRack slice-triggers, chopped audio fired from pads rather than played notes, excluded to their own bucket. And the standing dedup lesson (D105) is applied to the estimate before anyone gets excited: unique pairs are likely 10 to 30 times fewer than the pre-dedup count. Still thousands. The anchor exists, in his own archive.
The probe also produced a workflow finding that reframes how the corpus reads his work: tracks he names melody or chords are almost all chops. Of 2,602 clips on melody-named tracks, zero survive the chop filter; of 1,008 on chords-named tracks, three. His melodic ideas often live as chopped samples, not played MIDI, so anchor role-labeling must be content-based rather than name-based, and the chop bucket itself holds real taste signal for a later layer. The inventory stays local to his machine; the task stays parked per the sequencing guard.
QA sealed the refill: 28 samples, 28 of 28 hashes exact, sorted into four buckets before any generation touches them. The sorting carries one ruling worth recording: a sample of his own audio was ruled out of the milestone count by the owner himself, stricter than the recommendation on his desk. The fixed star stays pure by his own hand.
The conservative default does the quiet work here: the five provisional files stay out of the count until the provenance table proves them clean, not the other way around. With the seal in place the lane is fully open: exposure six is next, a drone-class sample through the complete fixed pipeline, true clock to ending-hang warn, published on the countersigned precheck’s green. Every fire from here builds the consecutive streak the milestone requires.
Owner-requested: yesterday’s adapter map, redrawn after the densest day of changes the pipeline has had. The brain is unchanged: one pretrained music transformer (360M parameters, never retrained whole) wearing thin LoRA adapters, each about 1% of the weights, each with one job, each traced to a verdict. What changed today is everything around it: how the sample is cut and clocked, how the engine decides to follow or invent, where the chords strike, and what stands between a generation and his ear. Hover over (or tap) any box in the flow for a plain-language note on what it does and why it exists.
The lineage: three adapters, one engine
Every adapter, its objective, its current standing
The ML flow, end to end, as of tonight
The full path a sample travels, rebuilt where today rebuilt it. Solid boxes run in production; IN WORKS is decided and being built; FUTURE is decided, not started; BANKED is parked with a written revival condition.
The reading that survived the day: the adapters carry who he is, the retrieval and routing layers carry what he plays, the screens carry what he rejects, and everything that broke today broke in none of those places: it broke in the plumbing that carries them to his ear. The engine that earned the 5/5/5 tonight is the same one that missed four times this morning. The difference was the chain around it.
Milestone datapoint two: 2 of 3 at fire, with a double 5. EH, a true drone (variance 0.071), B minor, never seen, through the same fixed pipeline that fired yesterday’s sweep. The clean era is now 2 for 2, the fires are consecutive, and the milestone gate needs three: exposure seven can close it.
The preregistration scored as well as the takes did. Boredom was named as the risk axis before the generator ran, and it was the only complaint class: zero timing complaints, zero wrongness. Both noted clips flagged the same thing, the first five-bar chord span holding too long, one of those notes sitting inside a 5. And his question on the takes became the next lever on the spot: all three takes shared the same one-chord-per-beat pocket, which grid-establishing forces by design, so the variety axis now extends to rhythm: vary the strike density and pattern per take while the piano keeps its clock role on weak-pulse ground.
Two system wins rode the datapoint. The precheck earned its first save: the gauntlet caught a held dissonance over the bass (the class his b42 ratings named scary, at double the warn line, in two of three clips) and stripped it before his ears; the stripped clip then took a 5, which vindicates the strip, removing the sting cost nothing. The selection-time version of that screen was built and unit-tested the same hour, so the post-fix retires. And the ablation from this afternoon’s external critique is cut as T048: the adapters-off control joins the interlude queue. EH burns; the counted pool stands at 12.
Owner-requested: the framing beyond batch ratings. Where this engine sits among music transformers, how much of the network is actually ours, and what can and cannot be claimed, in research terms and plain ones. This entry refreshes at era boundaries, not daily.
What we started from
Research terms: the base is the Anticipatory Music Transformer, Medium (Thickstun et al., 2023; arXiv 2306.08620), a 360M-parameter autoregressive event-token transformer with a 27,512-token arrival-time vocabulary, trained on the filtered Lakh MIDI corpus: 164,747 MIDI sequences, about 7,827 hours of music. The license, precisely: code and weights are Apache 2.0; the Lakh MIDI dataset is CC-BY 4.0, and many of its files are themselves transcriptions of copyrighted songs, a caveat Stanford's own model card states and the reason a clean-source model swap sits on this project's ship-time list. Its defining trick, anticipation, makes conditional infilling (write a melody given the chords) a native capability rather than a hack. In plain terms: a mid-sized music brain that read a large library of everyone’s MIDI, is permissively licensed for research, runs on a home graphics card, and was born knowing how to answer other music, which is the one skill this whole project depends on.
Where it sits
| generates | size | built for | personal? | |
|---|---|---|---|---|
| this engine | MIDI notes, into his DAW, over his sample | 360M + three ~1% adapters | one producer’s taste, sample-aware | entirely: the point |
| AMT base (our start) | MIDI notes, general | 360M | average competence over Lakh MIDI | no: sounds like everyone |
| MusicGen family (Meta) | finished audio from text prompts | ~300M to 3.3B | whole songs and textures from a description | style prompts, not a person’s corpus |
The comparison to hold onto: audio models make finished sound you cannot pull apart; this engine makes notes, editable in his DAW, locked to a sample he chose. Different species, and the species was chosen deliberately: per his standing rule, the project is MIDI-first. Prompting a big general model was never tested head-to-head, and the methodology page now says so plainly; the closest control on record is b25, where this base with his data beat the previous stock-trained model at his ear.
How much of it is ours
Research terms: the personalization is low-rank adaptation, rank 16, alpha 32, on attention and projection matrices, three adapters continue-trained in a lineage (cohesion, then his sound, then loved movement), each adopted by blind A/B at the ear, with held-out validation floors verified by independent re-descent (gold: 0.237, twice, 0.002 apart). In plain terms: we never rewired the brain. We hung three small steering weights on it, each about a hundredth of its size, trained on one man’s archive, and the measured result went from competent-but-generic threes to a 5/5/5 on a song it had never heard.
What cannot be claimed, and is not: no benchmark scores against other systems exist (they generate finished audio, we generate DAW-ready MIDI; the shared metric would be his ear, and his ear has only rated this engine), and the adapters-off ablation that would isolate the adapters’ exact contribution is queued but not run. When it runs, this entry gets its missing number. The deeper comparisons, this engine against the two-tower audio-conditioned student, belong to the endgame era.
The refill left five samples in the provisional bucket because nobody could say for certain they were absent from the taste-training sources. Two independent instruments now say they are. A two-pass name sweep came back clean, and a content sweep compared each file against all 1,377 cached embedding windows of the liked-tracks corpus: the maximum cosine similarity found was 0.889, which is style-neighbor territory, against a same-track signature of above 0.95 established by the backbone bake-off. The decision rule was fixed before the sweep: both instruments clean means a promotion recommendation, and the recommendation is now with QA and the team lead: promote all five to counted, which would put the milestone pool at 16 to 17.
The method note that makes this citable: the content instrument is the same embedding backbone whose take-ranking died at 0 of 6, used here strictly inside the scope its backtest left it (style-level similarity), and the conservative default held throughout: the files stayed out of the count until the evidence arrived, not the reverse.
Between verdicts, the owner sketched the VST’s product shape, now ledgered (D194): a stem-splitter mode that, when the user chooses it, generates the complement in separate parts (chords, melody, bass as their own stems) instead of one combined layer, a MIDI-track architecture so those generations land as editable tracks rather than a bounced blob, and retro MIDI drag-out, pulling any past generation out of the plugin into the project. The design note landed within the hour as a phased architecture, milestone lane untouched. It joins the endgame note as the second vision entry in one day that came from his usage instincts rather than a feature list: he described how he would actually reach for the thing while it was mid-test under his hands.
No fire: 3, 3, 2, one take short of the 2-of-3 gate. The clean era stands 2 for 3, the milestone streak resets, and the curve reads fire, fire, miss-with-mechanism. The ground was the clean era’s first pulsed fresh test: real drums in the window (onset pulse 5.54, against the 2.4 detector line), so the character dial gave its third distinct answer in three clean exposures: sparse ground got follow, the drone got grid-establishing, this got metric strikes on the beat.
The mechanism was verified before any theory was written. His first instinct, that something read as slightly detuned, was tested and cleared: the cut measures 1.0 cent flat, in tune. The real cause came from a harmonic-percussive decomposition of the window: every bar sounds the same six pitch classes constantly, a dense in-key wash, and the engine’s triad complements doubled notes an already-full bed was sounding. In key, so nothing was wrong; redundant, so nothing was added; and the doubled voicings against the sample’s timbre are the likely source of the detuned feeling. The routing was correct by its own rules, because the follow gate measures harmonic change (0.041 to 0.102 here, correctly free-gen), and change was never the problem. Density is the fourth character axis, exposed by this miss: crystalpeak taught follow-when-moving, this teaches thin-when-thick.
The lever was built the same hour, to his own spec, paraphrased: go to a higher register when the mids are already chorded out, or answer the sample’s chords in their gaps when they are not strong, with shorter plucks and the melody forward. Complement, in his vocabulary, means bounce and pocket with the sample first and frequency room second, now law L25; the one gap it exposes, straight-grid onsets with no swing matching, is banked as a designed lever awaiting its first ear evidence. Two branches (subtractive shell voicings lifted above the wash; call-and-response plucks reusing the gap machinery that won b39), calibrated so this ground routes to call-and-response while both fire batches route unchanged, regression clean. Span-v2’s own test came back orthogonal: zero too-long and zero too-changey complaints, so it stays unbanked and the next drone judges it. Two smaller notes for the record: the selection-time dissonance screen made its first production catch on this batch (a held flat-9 stripped at construction, one batch after the post-fix it replaced), and the five provisional files were promoted on the provenance evidence: 21 counted samples remain after this one burned.
The day closed with a full-picture review, and it belongs here in the same plain language it was given. Where we are: in about four days, from asking whether a model trained on his old projects could make anything decent, to a machine that generates parts on never-heard ground that he rated as usable. That is the headline, and the path has weak spots worth naming.
What holds up. The core bet is the moat: the taste comes from his projects and likes, the filters from his verdicts one at a time, and nobody can download that; the 5/5/5 was his taste generalizing, not a generic model getting lucky. The measurement system earned its keep twice today: sealed samples, predictions before listening, misses that stay misses, and seven wrong theories died safely inside that discipline instead of shipping as fixes. The debugging culture works because his casual comments are treated as data and tested mechanically; both era-defining bugs fell to that. And the mouth-to-code loop runs inside an hour, which is the project’s real engine.
What does not, yet. The milestone rate rests on three clean datapoints, which is barely a sample: the 2-of-3 could be luck in either direction, and the boring fix is ten-plus fresh samples before the rate is called real. The ground truth is one human, and humans are noisy: mood, fatigue, and a deliberately rising bar all move scores, so a quiet calibration is planned, re-serving an old batch to measure how much slack a verdict carries (the methodology page’s test-retest gap, finally getting its measurement). The tempo stamp is the humility lesson: the whole team built seven musical theories on a three-byte clerical error, and the standing question is what else like it is still in there. Complexity accrued four behaviors today alone, which is exactly how the lost-at-batch-200 fear comes true, so pruning reviews have to actually delete things. The code is research scaffolding, not a product, and today’s wins did not shrink that distance. And the melody ceiling still stands, corpus-capped, with his game-soundtrack reference pointing at the fix: a small hand-picked gold set of melodies he loves.
The order going forward: keep the fresh-ground lane hot until the fire rate means something (ten-plus samples); prune on schedule, deleting what evidence no longer supports; attack melody with the reference tracks plus a small gold set; harden the pipeline into one codebase only when the rate holds; and the self-farming endgame stays last, exactly as guarded. The short version the review ended on: the science is ahead of the engineering, the evidence is thinner than the excitement, both are normal for day four, and a machine that catches its own mistakes is worth more than any single 5/5/5.
The review became the standing order (his priorities replaced the team’s, verbatim in the ledger as D199), and the execution plan landed an hour later, committed as the module roadmap. One rule governs it: modules advance on exit criteria, never on dates or excitement, and every module carries a written kill clause. The order: prove it, then polish it, then productize it, then scale it.
| module | the work | exit / kill |
|---|---|---|
| 0 · measurement (now) | fresh exposures until the fire rate is a real number (10+ fair tests; 3 exist: fire, fire, miss); the pruning review; the blind re-rate calibration (T050, his own proposal) | kill: rate lands well under 2-of-3 → stop, diagnose which ground types miss |
| 1 · melody | the one quality axis with no dial: a small hand-picked gold set from his reference tracks, cooked onto the proven lineage | kill: degenerates like pure-likes did → bank the loss |
| 2 · hardening | collapse the three-environment, this-PC scaffolding into one codebase with a regression suite where every rated batch re-runs identically | exit: a stranger’s machine runs an exposure from a fresh checkout |
| 3 · product core | capture, generate, one role live in the DAW; the retro roll UI (v0 shipped in tonight’s rating app) and drag-out | exit: he uses it in a real session without being asked |
| 4 · arrangement | stem-splitter mode: up to four role generators complementing the sample and each other | exit: an arrangement he keeps |
| 5 · flywheel | the archive pairs, the stacks offer, the self-farming dataset | gated: only once the teacher is trustable; bias governors already written |
One methods note rode the plan discussion and belongs on the record: the owner’s own estimate of his rating wobble is high: by his read, a re-served sample could plausibly come back fire much of the time even where it previously missed. That estimate is exactly why the milestone counts only first-listen fresh ground, why single verdicts carry slack until n grows, and why T050’s sealed re-serve exists: to turn the wobble from a guess into a number every future verdict gets read against. The day ends where the team lead framed it: the machine learned to measure, and then the owner taught it to plan.
The milestone bar is now locked in the owner’s own units (D203): a sample passes when 2 of 3 regenerations fire. Two numbers get tracked separately from here: regen reliability per sample, and coverage across samples. On the locked definition the board reads 2 samples pass, 1 fails (the 135bpm loop and the drone pass; the dense wash fails), with sample four on his ears now. The wording of every future preregistration speaks this language.
Exposure eight carries three firsts. The rhythm-pattern library is engaged for the first time, one pattern per take (steady quarters as control, a busier grid, a half-time lean), which is his same-pocket question from the drone verdict answered at the ear, per-pattern predictions locked. The retro roll made its debut in the rating app: each take now draws its notes as a pixelated piano roll, melody red, chords navy, the first light of the product’s drag-surface. And the preregistration is written in the locked bar’s units. The vocabulary law also minted as L25: complementary means bounce and pocket first, frequency room second, with the one exposed gap (straight-grid onsets, no swing matching yet) banked as a designed lever awaiting ear evidence.
The product thread closed the evening (D204): the front panel, in his sketch, is a depth dial, a pocket control, a regenerate button, and the retro roll, governed by one law: every knob only ever selects behaviors with verdicts behind them, never an experiment. The regenerate button’s promise is the milestone bar. The evening in one line, the team lead’s: the machine measures while the owner designs, and neither waits on the other.
Module 0 is the standing work, and exposure eight is on his ears, sample four of the locked bar, on the last of tonight's calls. The gate is unchanged, three consecutive fresh fires, and the board reads fire, fire, miss-one-short with every miss mechanism measured. Aboard next time: the density-complement lever (thin shells or gap plucks on dense ground, built to his spec, fires regression-clean), span-v2 still unbanked pending a drone, the rhythm-pattern library riding the next weak-pulse ground, and T048’s adapters-off ablation in the interlude queue. Behind the verdict, packaged and held: the dense-wash re-run on the new density lever, whose lane now runs under a written six-governor protocol, the load-bearing one being that a fire on re-served ground banks nothing: the lever banks only when fresh dense ground confirms it. Pool: 21 counted. Entries land here as the day happens.
Sources