Notes & essays

Blog

Notes on reasoning, learning, and building AI systems.

From Off-Policy Data to On-Policy SFT

Can we keep the efficiency of SFT while reducing forgetting? We learn a sampler that brings expert data closer to the student’s policy, shifting repeated data-generation search into sampler training.

LaDiR: Latent Diffusion Enhances LLMs for Text Reasoning

Why use diffusion to model reasoning? The ideas behind LaDiR: generating thoughts in a learned latent space, training on the model’s own trajectories, and using denoising to control inference-time computation.