From Off-Policy Data to On-Policy SFT
Can we keep the efficiency of SFT while reducing forgetting? We learn a sampler that brings expert data closer to the student’s policy, shifting repeated data-generation search into sampler training.
Notes & essays
Notes on reasoning, learning, and building AI systems.
Can we keep the efficiency of SFT while reducing forgetting? We learn a sampler that brings expert data closer to the student’s policy, shifting repeated data-generation search into sampler training.
The format that teaches a model need not be the format in which it thinks. Uni-LaDiR learns intermediate thoughts from what later reasoning needs, and learns to generate them with diffusion.
Why use diffusion to model reasoning? The ideas behind LaDiR: generating thoughts in a learned latent space, training on the model’s own trajectories, and using denoising to control inference-time computation.