From Off-Policy Data to On-Policy SFT
Learning an amortized sampler for the information-constrained target distribution.
TL;DRQuick summary
- Problem: SFT is efficient, but fitting off-policy demonstrations can move a model far from its starting behavior and contribute to forgetting. Can we bring the data closer to the student while preserving the information it needs to learn?
- Our solution: Train an expert-conditioned sampler to match the information-constrained student distribution. A group-relative GFlowNet loss removes the learned normalizer; sampler training absorbs work otherwise repeated during data generation.
- Two uses: Offline, adapt an expert corpus for a chosen student. Online, refresh a LoRA sampler as that student learns, producing new targets for ordinary SFT.
- The real test: Better accuracy at matched total compute, while retaining prior capabilities. A learned sampler addresses repeated search and stale data; verifier errors, incomplete coverage, and forgetting still need separate evaluation.
- Draft estimates, not measurements: Offline about +2 percentage points and online +3–5 points over MCMC + SFT; no loss of prior-task accuracy is the retention target. These placeholders await experiments.
Project note · Method implemented · Our numerical results below are draft estimates
SFT’s off-policy mismatch
Supervised fine-tuning (SFT) is attractive because the training loop is simple and fast. Given a dataset of responses, we train the model to predict their tokens. Those tokens are already available, so their losses can be computed in parallel. Each update needs no fresh autoregressive rollouts, reward-based advantage estimates, or teacher scoring of newly generated responses—the extra work involved in on-policy RL and distillation. [2] [11]
But learning a new task can damage abilities the model already had. A model may improve at math while becoming worse at following instructions or answering general questions. This is catastrophic forgetting. Recent comparisons find that SFT often forgets more than on-policy RL, even at similar performance on the new task. [9] [10]
One reason is the distribution SFT asks the model to learn. Its responses usually come from a human or another model, rather than the student’s current policy. That makes the data off-policy. Even when every answer is correct, the demonstrations may use reasoning paths, wording, and intermediate steps that the student rarely generates. SFT asks it to reproduce that whole distribution.
Let \(q_{\mathrm{data}}(y\mid x)\) denote the demonstration distribution and \(p_\theta(y\mid x)\) the student. For a fixed prompt \(x\), ordinary SFT minimizes:
The entropy is constant, so the objective pulls the student toward the demonstrations. There is no term here that keeps it near its starting policy \(p_0\). If the model can represent the data distribution and fits the population loss exactly, its fitted distribution is \(p_{\mathrm{fit}}=q_{\mathrm{data}}\). Its distance from the starting model is therefore:
Fitting distant data means moving toward a distant policy. Some change is needed to learn the task; matching the expert’s particular way of solving it can demand more. Because the same parameters support many skills, that movement can disrupt prior behavior. RL’s Razor connects larger distribution shifts with greater forgetting and analyzes a minimum-KL bias of on-policy learning in an idealized policy class. Retaining by Doing provides complementary evidence: refreshing SFT data from the evolving student reduces forgetting in its experiments. The connection is useful, but training-task KL alone does not guarantee retention on other tasks. [9] [10]
On-policy methods address the mismatch by changing where training trajectories come from:
- RL, such as GRPO, samples responses from a recent student policy, scores them, and updates the student using their relative rewards. The feedback is attached to behavior the student actually produces. [11]
- On-policy distillation (OPD) also samples from the student, then asks a teacher to score its tokens. This directly addresses the mismatch between teacher-written trajectories and student-visited contexts; Thinking Machines also shows how it can recover instruction-following behavior after further training. [2]
Both approaches keep generation and feedback inside the training loop. We ask whether we can keep ordinary SFT as the student update and handle the mismatch in the data instead: can we transform off-policy demonstrations into a distribution closer to the student, preserve what they teach, and reduce forgetting?
Finetuning with Sampling pursues this route with MCMC. [1] Our project learns an amortized sampler for the same distribution-matching objective, replacing repeated per-example search with a reusable generation policy. The next step is to define precisely what “closer to the student” should mean.
What is the mismatch we want to remove?
Equation (1) tells us where SFT will try to move the student. We can choose that destination: replace the original demonstration distribution with one that preserves the required information while staying as close as possible to the student. This turns the motivation above into a data-distribution optimization problem.
For one prompt \(x\) and expert response \(\tau\), let \(C_\tau\) be the set of acceptable responses. In math, this can mean responses with the correct final answer; a stronger verifier can also check the reasoning. For the next few equations, write \(p(y)=p_\theta(y\mid x)\) for the fixed student and \(q(y)\) for the data distribution we are choosing. Our objective is:
The constraint says that the training data must preserve the expert information. The KL says that, within this constraint, its distribution should stay close to the student. Unlike the SFT objective in equation (1), we now hold the student fixed and change the data distribution. This is the information-projection objective used to formalize student-compatible data: choose the closest distribution among those satisfying the task constraint. [1] [9]
A natural candidate is to keep the student’s probabilities for acceptable responses and renormalize them. Let \(Z\) be their total probability:
Why is this the closest distribution? On the valid set, \(p(y)=Zp_C(y)\). Substituting this into equation (2), for any admissible \(q\), gives:
The second term is constant. The first is nonnegative and becomes zero when \(q=p_C\), which proves that equation (3) minimizes our objective.
If the original demonstrations already satisfy the same constraint, they are one feasible choice of \(q\). The optimum therefore obeys \(D_{\mathrm{KL}}(p_C|p)\leq D_{\mathrm{KL}}(q_{\mathrm{data}}|p)\). This is the precise improvement we seek in the data distribution. Whether SFT on those data also retains more prior capability is what our experiments must establish.
We now know exactly what “more on-policy” means here: generate from the student’s distribution conditioned on preserving the expert information. We discard unacceptable responses and retain the student’s relative preferences among the acceptable ones. If those responses are rare, even this best possible match can be far from the original student; the constraint still has to be satisfied.
Amortize the search into sampler training
Knowing the target distribution does not make it easy to sample. We could draw from the student and reject invalid responses, but that costs \(1/Z\) attempts per accepted response on average. This becomes impractical precisely when expert information is most useful: when the student rarely solves the problem on its own.
MCMC uses the expert trace to guide this search. For a given prompt, it proposes changes to the current response, scores them, and accepts or rejects each proposal. The response changes; the model weights stay fixed during this search. A new prompt requires another chain. [1]
Amortized inference moves much of this repeated search cost into training an inference machine. The reusable result is a set of sampler parameters: training on one batch can change how the sampler generates responses for later prompts. This is the training-for-inference tradeoff described by Bengio: invest computation in learning a reusable inference procedure, then use it to generate samples. [8]
In our case, the inference machine is a conditional sampler \(q_\phi(y\mid x,\tau)\). During fitting, it explores responses, receives student likelihoods and validity feedback, and learns to approximate \(p_C\) across expert examples. We train it directly from these scores, without MCMC-generated teaching examples.
At data-generation time, MCMC repeatedly revises a candidate response. The trained sampler instead starts with an empty response and appends tokens, using transition probabilities learned across examples. This constructive policy is the reusable object in a GFlowNet. [12] It still requires autoregressive decoding and verification. Both routes then use their verified outputs for SFT. Here, sampling-time inference refers to creating training data.
The intended benefits follow from this reuse: lower cost per generated example after training, shared learning across related prompts, and a sampler that can be refreshed as the student changes. The training investment only pays off if enough useful data is generated. Total cost must therefore include sampler training, student scoring, verification, and rejected samples.
One separation is essential. The sampler sees the expert response to help find acceptable outputs. The student scores each candidate using the prompt alone. Thus expert information defines what to preserve, while the student defines the density we want to learn.
From the KL target to a trainable loss
We can now derive the sampler’s training rule. Equation (4) says that minimizing the constrained KL is equivalent to matching \(p_C\). Define its unnormalized density as \(R(y)=p(y)\mathbf 1[y\in C_\tau]\). Our target is therefore \(q_\phi(y)=R(y)/Z\).
The problem has become distribution matching: we can score any candidate with \(R(y)\), but cannot enumerate all responses to compute \(Z\). The following steps turn that target into a loss we can evaluate on sampled responses.
Match probabilities up to one common scale
Assume the sampler can represent the target and assigns positive probability to its valid responses. For a valid response, rearrange the matching condition and take logs:
The unknown value on the right is the same for every valid response to this expert example. So the sampler is correct when every response has the same sampler-to-target log gap.
This is the trajectory-balance condition for an autoregressive GFlowNet. The connection does not require a different generator: the sampler still appends one token at a time, and its complete-response probability is the product of those token probabilities. Each response has one path through its prefixes. Matching its probability to \(R(y)/Z\) therefore matches the probability of the full generation trajectory. Our earlier work applies this flow-based view to language and visual reasoning. [3] [4]
Introduce a scalar \(z\) to represent the unknown \(\log Z\), and square the error in equation (5):
This is the GFlowNet trajectory-balance loss. It is computable from the sampler’s token log probabilities, the student’s score, and one common offset. Both too much and too little probability create an error. If the error vanishes across the valid set, \(q_\phi(y)=e^{-z}R(y)\); requiring the probabilities to sum to one forces \(e^z=Z\). We recover \(q_\phi=p_C\), the minimizer of our original KL.
We have replaced an intractable normalized target with a sample-level training error. The two objectives have the same ideal solution; reaching it still requires exploring the valid responses, not merely fitting a few observed ones. [6]
Eliminate the offset with a group of responses
Rather than learning \(z\) with another model, we can solve for it within each rollout group. Generate \(K\ge2\) valid candidates for the same prompt and expert demonstration. Write their log gaps as \(a_i(\phi)=\log q_\phi(y_i)-\log R(y_i)\), and let \(\bar a=K^{-1}\sum_i a_i(\phi)\) be the group mean before the update.
Averaging equation (6) gives \(K^{-1}\sum_i(z+a_i)^2\). Its derivative with respect to \(z\) is \(2(z+\bar a)\), which is zero at \(z=-\bar a\). In other words, the best common offset centers the group’s log gaps. Substitute this offset into equation (6), average over responses, and we obtain:
This is the group-relative matching loss: make the sampler-to-target log gap agree across responses. The operator \(\operatorname{sg}\) means that the measured group mean is held fixed during backpropagation. An above-average gap means a response is overrepresented relative to the group, so the loss calls for lowering its log probability. A below-average gap calls for the opposite adjustment. We learn the relative probabilities without learning the normalizer. [3] [4] [5]
Each expert example gets its own group mean because it has its own \(Z\). This mean is a batch offset, not an exact log-normalizer estimate before convergence. Equation (7) uses one update per fresh rollout group. Reusing samples for multiple updates additionally requires accounting for the changed sampling policy, for example with importance weighting. [6]
For a valid candidate, \(\log R(y_i)=\log p_\theta(y_i\mid x)\). Training therefore needs only sampler and student log probabilities: compute their difference, subtract the group mean, and minimize the squared residual. Use complete sequence log probabilities, including termination; length-normalized scores would define a different matching problem.
The derivation assumes a distribution supported on acceptable responses. In practice, the generator can produce invalid outputs, so we match its accepted-output distribution. Conditioning on acceptance subtracts the same log acceptance probability from every valid output’s log probability. That constant cancels in the group-centered residual, making equation (7) computable with the raw sampler probabilities.
This cancellation does not train away invalid mass. Verification determines what enters SFT, while acceptance rate and coverage must be evaluated separately. The sampling policy must also agree with the probabilities in the loss; changing temperature or truncating the rollout distribution requires corresponding correction.
We have arrived at a trainable sampler without changing the student’s SFT objective. Next, we can either fit that sampler once before SFT, or keep fitting it as the student evolves.
Offline: train the sampler, then run SFT
The offline recipe has three stages. The student does not change while we train the sampler.
Offline · fit once, then SFT
fit(sampler, fixed student) # Eq. 7
data ← verified_samples(sampler)
SFT(student, data)
fit updates only the sampler using fresh rollout groups and equation (7). Both fit and verified_samples condition generation on the prompt and expert demonstration; the latter retains only verified responses and holds the fitted sampler fixed. SFT updates only the student, using prompts and sampled responses; the expert demonstration is not an extra student input.
This version replaces per-example MCMC data creation with a trained, reusable generator. It pays an up-front sampler-training cost, then shares that work across examples. Whether the reuse saves compute is an experimental question: training, student scoring, verification, and failed generations all belong in the cost.
The sampler continues to target the starting checkpoint, even after SFT changes the student. That is a deliberate property of this offline recipe.
Online: a LoRA sampler on the evolving student
The offline sampler targets the starting checkpoint. Online training instead refreshes the data distribution after the student changes. We implement the sampler as a LoRA adapter on the current student backbone, rather than maintaining a separate full-size sampler model. [7]
For an adapted weight matrix, the sampler uses a low-rank update:
The same backbone has two roles. Adapter enabled: generate with the prompt and expert demonstration. Adapter disabled: score candidates under the student using only the prompt, or update the student with ordinary SFT. We keep the adapter separate from the final task model.
This design has three practical advantages. The sampler update trains a small set of parameters, reducing its additional parameter and optimizer-state storage. It starts from the student’s existing language and reasoning capabilities instead of learning a generator from scratch. And as SFT improves the shared backbone, that improvement is immediately available to the sampler. These are architectural reasons for LoRA; whether they reduce end-to-end compute or improve accuracy is measured separately.
Sharing the backbone also creates a dependency. Even with fixed adapter weights, updating \(\theta_t\) changes the sampler’s distribution. We therefore refit the adapter against the current constrained target:
Online · refresh as the student learns
sampler ← LoRA(student)
repeat:
fit(sampler, fixed student) # Eq. 7
data ← verified_samples(sampler)
SFT(student, data) # LoRA off
fit updates only the adapter, scoring candidates with LoRA disabled. The SFT phase disables the adapter and updates the backbone. Each new round fits the adapter against that updated backbone.
The expert demonstrations remain fixed; the generated SFT targets can evolve. Refreshing the adapter helps address stale data, but the refresh frequency has a cost. Small adapter updates also have limited capacity. Both the update schedule and LoRA rank therefore belong in the online ablation, rather than being treated as automatic improvements.
This differs from on-policy distillation, which samples student trajectories and uses teacher probabilities as feedback. [2] Our expert-conditioned adapter generates data for the student’s constrained distribution; the student receives ordinary SFT.
Offline experiments
Setup
The offline comparison follows the math setting of Finetuning with Sampling: Qwen2.5-3B, MATH levels 3–5, with 8,230 training problems and 1,024 test problems. We evaluate single-shot accuracy on MATH, AMC, MATH500, and GSM8K, plus Chemistry, MMLU, and GPQA for prior-capability retention. [1]
The source SFT search uses 1–2 epochs, learning rates {5e−5, 1e−5, 5e−6}, and batch sizes {16, 32, 64}, with AdamW and a cosine schedule. Its MCMC baseline uses 10 transitions, block size 32, and maximum sequence length 1,856. [1] Our offline sampler is fitted to the starting student, then frozen to create a dataset before SFT begins.
For these placeholders, each percentage comes from an illustrative integer correct count in one evaluation pass, then rounds to one decimal. The sizes are MATH 1,024; AMC 83; MATH500 500; GSM8K 1,320; Chemistry 600, following the paper and its evaluation files. For example, 25/83 rounds to 30.1% on AMC; MATH500 moves in 0.2-point increments.
MMLU uses the full 14,042-item test set and a micro-average. GPQA temporarily assumes the 198-item Diamond subset; the baseline’s variant needs confirmation before a measured comparison. The prior-task average is the unweighted mean of the three task scores.
Results: task accuracy and retention
Draft estimates, not experimental measurements. Baseline scores are reported in Table 1 of Finetuning with Sampling, converted to percentages. Rows marked Est. † are unmeasured planning values: gains vary around +2 points offline, with slight task-level variation in retention. Published baselines retain their original rounding.
Learning the new task
Reported baselines; offline estimates vary around a +2-point gain.
| Method | MATH | AMC | MATH500 | GSM8K | Δ MATH |
|---|---|---|---|---|---|
| Base model | 31.5 | 13.3 | 24.5 | 57.9 | −18.0 |
| Expert-data SFT | 24.3 | 10.0 | 16.8 | 45.5 | −25.2 |
| OPSD | 26.7 | 8.4 | 33.2 | 52.4 | −22.8 |
| GRPO | 45.7 | 24.9 | 31.3 | 80.8 | −3.8 |
| UFT | 47.0 | 29.3 | 29.7 | 74.6 | −2.5 |
| MCMC + SFT | 49.5 | 27.7 | 58.2 | 78.2 | +0.0 |
| Our offline sampler Est. † | 51.6† | 30.1† | 59.8† | 80.2† | +2.1† |
Δ MATH is the percentage-point change from MCMC + SFT. † Draft estimates; not measured.
Keep prior capabilities visible
Per-task changes reveal losses that an average can hide.
| Method | Chemistry | MMLU | GPQA | Prior avg. | Δ vs. base |
|---|---|---|---|---|---|
| Base model | 28.3+0.0 pp | 65.1+0.0 pp | 33.3+0.0 pp | 42.2 | +0.0 |
| Expert-data SFT | 22.2−6.1 pp | 64.8−0.3 pp | 29.8−3.5 pp | 38.9 | −3.3 |
| OPSD | 24.2−4.1 pp | 65.2+0.1 pp | 31.3−2.0 pp | 40.4 | −1.8 |
| GRPO | 27.8−0.5 pp | 65.2+0.1 pp | 31.3−2.0 pp | 41.4 | −0.8 |
| UFT | 28.3+0.0 pp | 65.3+0.2 pp | 32.8−0.5 pp | 42.1 | −0.1 |
| MCMC + SFT | 26.6−1.7 pp | 65.1+0.0 pp | 34.3+1.0 pp | 42.0 | −0.2 |
| Our offline sampler Est. † | 28.3†+0.0 pp | 65.0†−0.1 pp | 33.8†+0.5 pp | 42.4† | +0.2† |
Small numbers show each task’s change from the base model. Δ uses displayed rounded scores. † Retention estimates; not measured.
The offline placeholders give 51.6% MATH, versus 49.5% for MCMC + SFT. Gains across the four math tasks range from 1.6 to 2.4 points. The prior-task average is 42.4%, but MMLU still slips by 0.1 point in this scenario. Learning the new task and retaining prior skills must be checked separately.
Forgetting needs a per-task check. The reported MCMC baseline is only 0.2 points below the base model on the prior-task average, yet Chemistry drops from 28.3% to 26.6%. An average can hide that loss. We therefore report all three prior tasks and their change from the base checkpoint. [1] A supported retention claim requires repeated runs and per-task confidence intervals with a prespecified tolerance for degradation; the estimated row is a target, not evidence of no forgetting.
For efficiency, compare equal-size verified datasets and include sampler fitting, generation, student scoring, verification, rejected candidates, and SFT in total compute. A one-pass rewrite and rejection-sampling SFT are useful controls: they test whether learning a distribution buys more than cheaper data generation alone.
Conclusion
The offline hypothesis is concrete: a reusable data-preparation model should preserve MCMC’s learning benefit while reducing the cost of producing enough training data. The draft accuracy target is roughly +2 points while keeping each prior capability close to its starting level. It succeeds as an efficiency method only if those gains survive a comparison at matched total compute. For a small dataset, sampler training may cost more than the search it replaces.
Online experiments
Setup
Use the same starting checkpoint, expert split, and evaluation suite. The sampler is a LoRA adapter on the evolving student. Compare three schedules at matched total compute: a sampler fitted to the initial checkpoint, an adapter left frozen while the backbone changes, and an adapter refreshed between SFT updates. The second control matters because a frozen adapter still changes its outputs when its backbone changes.
Report LoRA rank, adapted modules, rollout group size, refresh interval, and both learning rates with the runs. Include MCMC sampling + RL as a stronger post-training comparator, in addition to MCMC + SFT. Its published scores are available in the same math setting. [1]
Results: projected online improvement
Draft estimates. Gains vary from 3.5 to 4.8 points over MCMC + SFT, rather than adding one constant to every task. Retention values include small gains and losses near the base checkpoint. All † entries remain unmeasured.
Does refreshing help?
Online estimates vary by task (+3–5 points); include the stronger RL pipeline.
| Method | MATH | AMC | MATH500 | GSM8K | Δ MATH |
|---|---|---|---|---|---|
| MCMC + SFT | 49.5 | 27.7 | 58.2 | 78.2 | +0.0 |
| MCMC + SFT + RL | 54.5 | 24.1 | 65.2 | 83.0 | +5.0 |
| Our offline sampler Est. † | 51.6† | 30.1† | 59.8† | 80.2† | +2.1† |
| Our online sampler Est. † | 53.7† | 32.5† | 62.2† | 81.7† | +4.2† |
Δ MATH is the percentage-point change from MCMC + SFT. † Draft estimates; not measured.
Retention through the online loop
Illustrative task-level variation near the base checkpoint, with losses shown explicitly.
| Method | Chemistry | MMLU | GPQA | Prior avg. | Δ vs. base |
|---|---|---|---|---|---|
| Base model | 28.3+0.0 pp | 65.1+0.0 pp | 33.3+0.0 pp | 42.2 | +0.0 |
| MCMC + SFT | 26.6−1.7 pp | 65.1+0.0 pp | 34.3+1.0 pp | 42.0 | −0.2 |
| MCMC + SFT + RL | 28.5+0.2 pp | 65.2+0.1 pp | 35.4+2.1 pp | 43.0 | +0.8 |
| Our offline sampler Est. † | 28.3†+0.0 pp | 65.0†−0.1 pp | 33.8†+0.5 pp | 42.4† | +0.2† |
| Our online sampler Est. † | 28.2†−0.1 pp | 65.2†+0.1 pp | 33.3†+0.0 pp | 42.2† | +0.0† |
Small numbers show each task’s change from the base model. Δ uses displayed rounded scores. † Retention estimates; not measured.
The online placeholder is 53.7% MATH, versus 51.6% offline: a 4.2-point gain over MCMC + SFT. It remains below the published 54.5% MCMC + SFT + RL result. The retention average is 42.2%, yet Chemistry is 0.1 point lower than the base. Neither task gains nor a stable average establish superiority over the stronger pipeline or prove an absence of forgetting. Prior-task retention must be checked at each checkpoint, because several individually small updates can accumulate into forgetting.
What would explain an online gain?
Three comparisons can turn an accuracy difference into an explanation:
- Tracking the student. Refreshing should help most when the student has moved away from the offline target. Compare refreshed and frozen samplers at equal compute, measuring downstream accuracy, accepted-output log-gap dispersion, and validity together.
- Reusing the adapter. Keeping the previous LoRA may reduce the fitting work needed after a student update. Compare retained and reset adapters against the same backbone; record updates and compute to reach comparable sampling quality.
- Refreshing at the right frequency. Frequent fitting may reduce mismatch while leaving less budget for SFT. Sweep the refresh interval under a fixed total budget, and track both new-task accuracy and prior-task retention.
These are proposed explanations to test, not observations from completed runs.
Conclusion
The online method treats data generation as part of post-training: the student changes, so the sampler learns a new target. The projected +3–5-point gain motivates the experiment, but the decisive result is an accuracy–retention–compute improvement over a fixed sampler and the stronger MCMC + SFT + RL pipeline. LoRA makes repeated fitting practical in parameter and optimizer storage; it does not by itself guarantee lower runtime or prevent forgetting.
Where this is useful
Offline, this is student-aware data preparation. An existing expert corpus can be rewritten into verified responses that better match a chosen model before SFT begins. The useful operation is more specific than removing bad examples: preserve what makes an example worth learning, while changing how that information is expressed for the learner. A fitted sampler can then process more examples from the same domain. Moving to a different student or a substantially different domain may require refitting.
Online, the same mechanism becomes a post-training method. A LoRA sampler keeps generating expert-informed supervision as the student evolves. This is useful when a fixed synthetic dataset becomes stale, or when on-policy exploration rarely discovers successful responses without expert help. The data distribution adapts; the student’s update remains ordinary SFT.
What does learning the sampler actually fix?
Our central argument concerns two costs: searching again for every example, and rebuilding data when the learner changes. Amortization can reduce the first through reuse; the online schedule addresses the second through refreshes. The paper’s repeated MCMC transitions and fixed-reference preprocessing motivate these questions. [1] Our assessment is that the project should be judged on those operational benefits, not merely on replacing one sampling algorithm with another.
There are three boundaries to that argument. First, at exact matching, MCMC and our offline sampler produce the same target distribution. An offline accuracy gain must therefore come from better approximation, coverage, or compute allocation in a finite-budget run. Second, a correct final answer does not establish that every reasoning step is sound; a learned sampler inherits the limitations of its verifier. Third, matching responses on the training prompts places no direct constraint on unrelated evaluation tasks. Preserving prior capabilities remains an empirical requirement.
This suggests a useful role for the project in the post-training stack: learn how to prepare data for the model, and reuse that preparation as the model learns. Offline, the output is a student-adapted training corpus. Online, it is a continually refreshed source of SFT targets. The value comes from making expert information economical to reuse without sacrificing its content or the student’s existing abilities.
References
[1] Aayush Karan, Sitan Chen, and Yilun Du. Finetuning with Sampling: SFT Learns Better Than You Think. arXiv:2610.02140v1, 2026.
[2] Kevin Lu and Thinking Machines Lab. On-Policy Distillation. Thinking Machines Lab: Connectionism, 2025.
[3] Fangxu Yu, Lai Jiang, Haoqiang Kang, Shibo Hao, and Lianhui Qin. Flow of Reasoning: Training LLMs for Divergent Reasoning with Minimal Examples. ICML, 2025.
[4] Haoqiang Kang, Enna Sachdeva, Piyush Gupta, Sangjae Bae, and Kwonjoon Lee. GFlowVLM: Enhancing Multi-step Reasoning in Vision-Language Models with Generative Flow Networks. CVPR, 2025.
[5] Xiaodong Liu et al. GFlowRL: Scaling Distribution-Matching RL to Large Language Models. arXiv:2607.13394v1, 2026.
[6] Xuekai Zhu et al. FlowRL: Matching Reward Distributions for LLM Reasoning. arXiv:2509.15207v3, 2025.
[7] Edward J. Hu et al. LoRA: Low-Rank Adaptation of Large Language Models. ICLR, 2022.
[8] Yoshua Bengio, with Edward J. Hu. Scaling in the service of reasoning & model-based ML. 2023.
[9] Idan Shenfeld, Jyothish Pari, and Pulkit Agrawal. RL’s Razor: Why Online Reinforcement Learning Forgets Less. arXiv:2509.04259v1, 2025. See §4–5 and Appendix A for the KL analysis and its assumptions.
[10] Howard Chen, Noam Razin, Karthik Narasimhan, and Danqi Chen. Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting. ICML, 2026. See §3–4 for distributional analysis and approximately on-policy SFT; Appendix A.5 discusses limits of KL as a predictor.
[11] Zhihong Shao et al. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. arXiv:2402.03300, 2024. See §4.1 for GRPO.
[12] Yoshua Bengio. Generative Flow Networks. 2022. Discusses learning sequential construction policies and contrasts them with MCMC sampling.
Citation
Please cite this post as:
Murray Kang. “From Off-Policy Data to On-Policy SFT.” October 2026.
@misc{kang2026onpolicysft,
author = {Kang, Murray},
title = {{From Off-Policy Data to On-Policy SFT}},
year = {2026},
month = oct,
howpublished = {Research blog},
url = {https://mk322.github.io/blog/learned-on-policy-sampler/}
}