Steering What Models Learn: Live, Concept-Level Control of Training Runs
See what a model learns as it trains, then steer it at the level of concepts and behaviors. First milestone: 4–6 weeks on fine-tuning.
Why mid-run is blind
Long training runs are mostly launch and hope. The controls we have mid-run are coarse: learning rate, data mix, rollback. None of them let a researcher say "learn this, not that." We shape models almost entirely before training (data, objective) and inspect them after. Between those two points, over thousands of GPU-hours, what the model learns is largely unobserved and uncontrolled.
A debugger for training
The idea is a tool that lets a human see what a model is learning as it trains and steer it with precision, at the level of concepts and behaviors rather than scalar knobs:
- Watch. Live probes, sparse-autoencoder features, and small capability evals, tracked over steps.
- Breakpoints. Pause or alert when a feature emerges, or when loss on a slice rises.
- Hot-reload. Change data weighting, loss terms, or gradient hooks without restarting.
- Branch. Fork from a checkpoint to test an intervention before committing to it.
- History. Every intervention is logged as a replayable diff, so the run stays reproducible and auditable.
Controls, from coarse to precise
Reweight examples by concept (e.g. everything that activates feature X)
Hot-add a penalty on a probe's output
Project a concept direction out of updates; gradient routing to localize knowledge
Preventative activation steering; concept ablation during training (CAFT)
The gradient- and representation-level methods already have early evidence in fine-tuning (gradient routing, Cloud et al. 2024; persona-vector preventative activation steering, 2025; CAFT, 2025). Nobody has yet packaged them as live, composable controls.
First milestone
In 4–6 weeks: a small transformer (~100M parameters) fine-tuned on a task with a known spurious shortcut. A user watches the shortcut feature form, suppresses it live with a gradient or representation intervention, and ends with a model that generalizes correctly out of distribution.
Baselines are no intervention, data-only reweighting, and post-hoc fixes after training. Metrics: out-of-distribution accuracy, in-distribution accuracy retained, compute overhead, and number of interventions needed.
Scope choice: start with fine-tuning, which is short, cheap, and has proven methods. Move to long pretraining once the control primitives and branching harness are solid.
Why it matters
Efficiency: fewer wasted runs, and failures caught early. Control: direct shaping of what models learn, rather than only filtering data. Safety and audit: a replayable record of how a model was shaped, and a way to keep unwanted traits from forming at all instead of removing them later.
The open questions are real. How does a human express "learn the task, not the shortcut"? How early do mid-run signals predict final behavior? How do we measure collateral damage when concepts are entangled? Which interventions help now but hurt the finished model? And how do we hot-swap hooks across a distributed job cheaply and deterministically?
Risks too: observation tools may miss what matters; interventions could hide a behavior rather than prevent it (held-out evals the tool never touches are essential); overhead at scale may limit this to fine-tuning for a while. That is the work.
