Early accessis open. Apply now.

Using interpretability to develop better LLMs and ML systems.

A short note on the direction we are pursuing at Aquin Labs, what we have built so far, and the research questions that follow.

Problem: Software engineering is built on reverse engineering and intentional control: you can inspect a system, localize a failure, change a part, and reason about what will happen next. Frontier ML mostly is not. We train opaque models, watch loss and evals, and when behavior is wrong we add data, ablate hyperparameters, or ship anyway. The gap is not only scientific curiosity. It is that we still cannot develop models with anything like the deterministic grip we expect from serious software: see inside, attribute cause, intervene, verify. That matters for frontier LLMs, and for any ML system where wrong behavior is expensive.

Direction: Reverse engineer intelligence in order to build and develop with intention. Mechanistic interpretability is one of the best current instruments for that: sparse features, causal attribution, steering, checkpoint diffs. The research program is to put those instruments into the development loop itself, not only into post-hoc papers. Loss and evals tell you that something is wrong or right; internals are a candidate channel for why, and sometimes for what to change next. Same thesis across frontier models, retrieval systems, scientific predictors, and safety-critical deployments: different probes, not a different idea.

Concretely: before spend, can we tell whether data will teach the intended structure or mostly shortcuts? During and after a run, can we see what actually moved relative to a base? On failure, can we localize behavior in the network or the data and intervene instead of only retraining blind? At release, can internal structure and behavioral probes agree enough to change a promote / hold call? How far that helps, and for which model classes, is still empirical. We do not treat it as solved.

What exists today: An end-to-end pipeline, via a CLI, to inspect a model, simulate a train before you burn GPU, watch a live run, diff checkpoints, steer and probe behavior, and evaluate before promote. It supports LLMs and embedding models, optionally with an SAE. Coverage includes attribution (trace, circuits, logit lens), SAE capture / train / align, weight and feature diffs, residual drift, trajectory, analytical and contrastive simulation, training-watch ingest, consistency / boundary / custom evals, red-team, and deception feature ranking. Results track locally and can land in a web inbox. Primitives for the direction above, not a finished development environment.

Research questions: Where in the development pipeline can reverse-engineering a model actually improve outcomes, rather than only describe a finished checkpoint? Can internal forecasts usefully complement another full train when choosing data or recipe? Do structure-level diffs catch failure modes that loss and standard evals miss? When does localizing a behavior make intervention better than collecting more data? Under what conditions do behavioral probes and internal structure agree enough to change a ship decision? And which of these transfer beyond LLMs?

Links: Docs, Research, Early access, Careers, aquin@aquin.app.

Work with us

Interpretability tooling, custom SAE databases, mechanistic audits, circuit reports, and hands-on research, experiments, and studies for teams of all sizes. Reach us at aquin@aquin.app

Not sure if Aquin is right for you?

Aquin