World Models
A collection of Aquin work on world models: learned dynamics, checkable environments, and interventions you can score against ground truth. Training runs use aq.
01 H3 Avatar
A talking-avatar LoRA on MiniMax-H3 Ref2VA: one character image plus dialogue becomes a clip with model-generated speech, lip sync, and light motion. The recipe is locked behind a local UI so users pick a reference, a line, direction, and a background — not checkpoints or audio modes.

Base model is MiniMax-H3 Ref2VA with a rank-16 / alpha-32 adapter on to_qkv, to_out.0, linear_1, and linear_2 (~10.7M trainable params). The locked checkpoint is h3-avatar-audio-v2-step1000: 481 accepted transcript-conditioned clips, 385 train, 1,000 optimizer steps at 384×384 with a 56-frame cache — about 2.6 passes over the train set.
Why transcript conditioning
An earlier audio pass improved lip timing but not reliable English: the data only carried a language tag, not the spoken words. Audio-v2 runs ASR, keeps short English phrases with usable confidence, and puts the transcript into the prompt so the model is conditioned on what the character should say. Step 1000 was locked before further joint video/audio training could drift already-good visuals.
Training and eval live in the repo with an AQ recipe handoff. Serving keeps the H3 base hot; the adapter itself is small. A Modal warm endpoint is the reference path; production should use a dedicated inference host.
02 Mario World Model
An action-conditioned JEPA world model for Super Mario Bros, with LeJEPA/SIGReg in latent space and probe gates that catch planning failures early. Trained with aq as an interpretability-in-the-loop run: a model that predicts perfectly can still fail to plan, and the gate stops the train before that waste lands.

A 9.8M-parameter JEPA on 8,000 episodes of World 1-1 (4.88M observations), predicting five steps ahead in a 192-dimensional latent. Mario is the vehicle; the reusable gate harness is the product. Prior work (LeMario) only found that height was missing after planning failed (y-probe R² = 0.188). Here the same failure is caught automatically at epoch 2 when any probe falls under threshold.
What SIGReg needs in this setting
SIGReg transfers out of image self-supervised learning into an action-conditioned video world model, but not unmodified. Report T / null(n), not raw SIGReg: under the Gaussian null the statistic floors at 0.51914 / n, so λ is not comparable across batch sizes without that ratio. Collapse-detection power scales with batch size, while video memory pushes the batch down. A trailing BatchNorm cuts the collapse signal about 25×. Temporally correlated windows break the iid null unless the statistic is scored per frame position. And in this setting the stop-gradient was still required: collapse remains the prediction objective's global optimum unless the target path cannot be chased.
Anti-collapse is not usable state
Preventing collapse did not preserve the state a planner needs. A SIGReg-only model stayed non-degenerate and still could not report Mario's height. An isotropic-Gaussian embedding is a floor, not a sufficient condition, which is why the run gates on a probe of control-relevant state per checkpoint.
With aux heads, in-distribution height probe R² reaches 0.943 against LeMario's 0.188 on about 4% of the planned schedule. Numbers above are from a 4,000-step run on one level; the aux column is trained to make height decodable, so it shows a cheap head fixes the problem, not that the self-supervised objective alone learned a better representation. Full write-up and limitations live in the repo docs.
03 Pixel-Art-GWM
A playable 21 MB action-conditioned world model of a 2D pixel-art character, trained from scratch on self-generated data with aq, and a measured failure mode when the agent is about 1% of the frame. 4.4M parameters, 17.7 MB on disk, about 100 fps on a laptop GPU.

The finding
Reconstruction objectives averaged over a frame are not neutral about what they preserve. When the controllable agent occupies 1.3% of the observation, such an objective can be reduced by discarding it, and standard metrics improve while it happens. A reconstruction containing no agent at all still scores 0.987 whole-frame accuracy.
Two changes fix it: spatially structured latents (the agent occupies specific cells and cannot be marginalised) and per-pixel classification over the known 21-colour palette (cross-entropy cannot blur; class weights can express relevance).
Results
Dataset: 2,025,385 frames across 35,000 episodes, generated rather than scraped. Released with the model and a full report.
