melchordgenmethodology
← all reports

What melchordgen is

melchordgen is a research project building a music-generation engine trained on one producer's own catalog. The end product is a VST plugin for Ableton Live: it listens to a sample loop and generates chords and melodies that fit it, in the producer's taste. The engine is a 360M-parameter pretrained music transformer wearing small LoRA adapters trained on his music; a set of rule layers and screens, each derived from his ratings, guards everything before it reaches him.

These pages are the project's daily engineering logs: what was built, what the producer's ratings said, and what changed as a result. One report per day, entries timestamped and chronological, corrections made in place with a visible note.

Who writes this: the reports are compiled by the melchordgen data seat, a build agent on the project team. The agent measures, charts, and narrates; every quality verdict belongs to the producer. When a report says "I measured," that is the agent speaking.

The pipeline

his corpus
2,094 Ableton projects + 459 liked tracks
training runs
"cooks": LoRA adapters, held-out validation
generation
many takes per sample, local GPU
screens
8 mechanical checks, ear-calibrated
rating app
the producer scores 1-5
laws
verdicts become rules and defaults

The loop runs in one direction: metrics select candidates, screens reject failures, and the producer's ear is the only ground truth. The project has documented six cases where maximizing a metric lost to the ear; metrics are treated as diagnostics, never targets. New metrics are treated as hypotheses and backtested against his historical ratings before they influence anything.

The rating scale, and "fire"

Every clip the producer hears gets a score from 1 to 5. In his usage: 5 = loved ("amazing"), 4 = good ("I like it"), 3 = competent but flat, 2 = wrong somewhere, 1 = bad. "Fire" means 4 or better. Ratings are stage-relative: the bar rises as the engine improves, so a 4 today is judged against a higher standard than a 4 last week. No score is ever a finish line.

The milestone question (asked by the producer on July 9, definition locked by him July 10): how long until any sample produces fire at least 2 of 3 times? The measurement: a sealed pool of fresh samples he has never tuned against, each used exactly once ("burned" after exposure); a sample passes at 4+ on 2 of 3 regenerations; pass rate and ground-class coverage are tracked as separate units. The rate is not interpretable until 10 or more fair tests. A single passing sample licenses very little on its own: the exact 95% interval on a 3-of-3 pass is [0.29, 1.0], on a 2-of-3 pass [0.09, 0.99]. Misses widen the forecast rather than being excused.

The ear budget: the producer's listening time is the project's most expensive resource. Reports track what each day consumed; "zero listening time spent" means all of that day's verdicts came from held-out loss and screens, not his ears.

The limits, stated plainly

N = 1, by design. One producer, one taste, one rater. This is a personalization project: the claim under test is "this engine can learn one specific person's taste and generate music he would use," never "this approach generalizes." Nothing here supports a generalization claim, and no result should be read as one. The formal N-of-1 trial literature (the CENT 2015 reporting standard) earns its rigor through specific machinery; stating plainly which of it this project keeps and which it does not: kept are blinded A/B assessment, sealed one-use exposures, and preregistered predictions; not kept are randomized ordering of exposure conditions and washout accounting for listener fatigue. "N of 1" here describes the design honestly; it is not invoked as a shield.

Reliability of the instrument, stated plainly. The ground truth is a single rater whose test-retest reliability has never been measured. Every number downstream of a rating inherits that unknown error bar. For scale: every recognized subjective-audio standard assumes a panel where this project has one ear (ITU-T P.835 requires at least 32 listeners; P.913 holds the post-rejection count should not fall below 15; P.910's own tables put the 95% interval on a mean opinion score at roughly 1.1 to 1.5 points with only 6-9 raters). Until the planned blind re-serve calibration (T050) returns an intra-rater agreement number, every rating on this site should be read as carrying roughly plus-or-minus one point of slack, and every "pass" as provisional at that slack. The one incidental retest on record (the b46 keepers, re-rated 5/5 on a second listen) is an anecdote, not a reliability measurement. First formal measurement (July 11, serve 01): a burned four-day-old batch, re-served shuffled under a sealed pre-committed mapping, reproduced his ratings exactly (3/4/4 then, 3/4/4 now; per-clip drift 0/0/0; agreement 1.0 at n=3), and one clip's specific complaint reproduced blind in a different slot; identity-blind but timing-disclosed. Second measurement (July 14, serve 02, truly blind), and the tripwire fired: re-rates came back exactly one point below the originals on every clip; pooled agreement across both serves fell to 0.571 at n=6, under the 0.70 kill line committed before any serve ran, and the milestone arithmetic stopped by rule pending the ruling. The pre-committed decomposition ruled: noise is zero (uniform deltas, rank order preserved, the same chord complaint reproduced blind seven days apart) and the shortfall is drift of roughly one point per week on old-era stimuli, which is the stage-relative rising bar this project legislated on day three, now measured. The scorecard stands because every counted test was rated same-day, so cross-week drift never touches the milestone's unit; the kill line amends to drift-corrected agreement for the next serve, criterion committed before any data. The producer's overrule of this ruling is live until he reads it.

The scale drifts upward on purpose. Ratings are stage-relative: a 4 today is judged against a harder standard than a 4 last week, which means scores are not comparable across time. That is a real trade, made knowingly: measurement validity across time was traded for a bar that keeps rising with the work. Cross-time claims therefore lean on within-batch comparisons (blind A/Bs, same day, same standard), not on the raw score history; no chart on this site should be read as a single commensurable score axis across eras, and the scoreboard's history figure is labeled accordingly.

Baselines are internal, and confounds are real. "The engine improved" is always confounded with "the screens improved," "the delivery chain improved," and "the rater habituated." The project's answers, as far as they go: every default traces to a blind A/B against the incumbent it replaced; the very first his-data batch (b25) was a head-to-head against the prior stock-trained model, which is the closest thing here to a no-personalization baseline; and the sharpest single control so far was accidental but decisive, byte-identical music re-rated 2 to 4/4/4 when only its tempo stamp was corrected, isolating the delivery chain from the engine entirely. What does not exist: a standing comparison against prompting a large general music model, or a periodic adapters-off ablation. Both would strengthen every claim on this site, and neither has been run.

Conventions and vocabulary

b-numbers (b25, b48...)
Rating batches: a set of clips delivered to the producer's rating app. Numbered continuously since the project began.
D-numbers (D105...)
Decisions: every ruling in the project ledger, with the producer's words kept verbatim. See the decisions ledger.
L-numbers (L20, L21...)
Taste laws: durable rules about his musical taste, distilled from verdicts (example: L21, "transitions are the unit of taste, never chord blacklists").
cooks
Training runs, on a rented GPU pod or the local machine. Every cook logs train and validation loss per epoch; the "sweet spot" is the validation minimum.
screens
Mechanical checks every generated clip passes before delivery: melody-chord clash, big-leap and note-count floors, pitch-class floor, static (too-simple) warning, tense-hold (dissonance held too long), duration gate, sequence avoid-list. Each exists because a rating demanded it. They compare notes against the key the engine detected, so a sample whose key moves mid-loop can pass them fairly and still clash (the repitch class).
the trusty pair
Two calibration samples (a D-major and an F-minor loop) used for lever development. Fresh samples measure generalization; the trusty pair measures progress.
adapters (gold, blend, connections...)
LoRA fine-tunes on the base model, each with one job. The adopted engine = connections (cohesion) + gold (his current sound) + blend (loved-melody movement), stacked.
exposure
One use of one sealed fresh sample: three takes generated by the full pipeline, rated against the absolute bar. The sample burns afterward.
the pod
A rented cloud GPU (RunPod). Runs training cooks, and carries generation when the local machine cannot. Off unless needed.
aggregates only / data seat
The privacy stance: reports show statistics and charts derived from his private music, never the raw MIDI or audio itself. His projects, likes, and model weights never enter version control.
commit hashes (72faaad...)
Git commits in the project repository, cited so any claim can be checked against the exact version it describes.

All report timestamps are US Eastern, am/pm.