DuoOPD: Learning from
Joint Teacher–Student Outcomes
for Multi-Task On-Policy Distillation

Ao Yu1Weibo Gao2Heng Zhou3Linan Yue4Rui Li1Suyi Liu1Yu Yan1Yizhong Zhang1Qi Liu1,†

1 University of Science and Technology of China2 The Hong Kong Polytechnic University
3 The University of Hong Kong4 Southeast University

† Corresponding author  ·  yuao@mail.ustc.edu.cn

Abstract

On-policy distillation (OPD) trains a student on its own responses with token-level feedback from a stronger teacher, yet the teacher can fail on questions the student already answers correctly, and how often each model succeeds varies across tasks. OPD ignores these outcomes and, on average, pushes down even the student's correct responses; gating feedback by student correctness fixes the direction but uses the teacher in the same way whether or not it succeeded. We introduce DuoOPD, in which the student's outcome sets the direction of feedback and the joint teacher–student outcome decides how the teacher supports it: when only the teacher succeeds, its verified answer becomes context for scoring the student's failed response, and when only the student succeeds, a weight shared within the task reinforces the whole response. A single rule covers all four outcome combinations without task-specific settings. Across Qwen3 and Llama, DuoOPD outperforms all five baselines in mean macro accuracy, improving over OPD by 2.58 and 5.98 percentage points, and it also leads on two further task mixtures spanning scientific calculation, instruction following, and code generation. Ablations show that outcome-based direction alone stays near the gated baseline, while the joint-outcome designs supply most of the gain.

Method

In on-policy distillation, the student generates an answer and the teacher scores its tokens. The teacher's preferences may still favor an incorrect response or give little support to a correct one. DuoOPD uses verified outcomes to decide whether each student response should be reinforced or suppressed.

When the teacher answers correctly and the student does not, the teacher receives its own verified answer as extra context while scoring the student's response. This helps direct correction along the student's existing answer. The student never sees the reference. When only the student is correct, DuoOPD reinforces its whole response with a positive weight shared among such successes within the same task and rollout batch.

When both models are correct, every student token receives positive feedback; when both are wrong, every token receives negative feedback. In these two cases, teacher preferences determine the strength of the feedback. The same four-outcome rule applies across tasks, with the teacher frozen throughout training.

Method overview comparing outcome-agnostic OPD with DuoOPD’s four outcome-dependent feedback rules.
Student verification sets the feedback sign; joint verification selects reference-conditioned scoring, task-local weight sharing, or original-context teacher preferences. Full formulation in the paper

Results

Test accuracy (%) is measured by avg@8: mean success across eight sampled answers per question. Trained methods are evaluated after update 60 and averaged over three training runs. Macro gives equal weight to the three tasks in each setting.

Bold and underlined values mark the best and second-best distillation methods in each column. Initial students and teachers are included as references.

Table 1. Biology, chemistry, and physics across two model families.
MethodQwen34B → 0.6BLlama3.1-8B-Instruct → 3.2-3B-Instruct
BiologyChemistryPhysicsMacroBiologyChemistryPhysicsMacro
Student (initial)37.3838.8833.3836.5436.6342.4444.9441.33
Teacher79.5681.3876.8879.2773.2575.0669.8872.73
OPD67.3867.4258.3864.3969.1970.9061.7567.28
EOPD68.0668.4460.9265.8168.7369.5461.2166.49
OPDVR66.9266.6960.9064.8369.7572.4063.9068.68
ExOPD67.5466.7958.9464.4268.3369.8860.8166.34
FiRe-OPD67.0067.2359.5064.5869.5071.8161.6067.64
DuoOPD (ours)68.6070.1962.1066.9774.4276.3369.0273.26

DuoOPD improves over OPD in biology, chemistry, and physics for both model families. The largest gains are in physics: 3.73 percentage points for Qwen3 and 7.27 for Llama. Averaged across the three tasks, the gains are 2.58 and 5.98 points, respectively.

Table 2. Additional task mixtures, each trained with Qwen3-4B → Qwen3-0.6B.
MethodScienceKnowledge, understanding, calculationHeterogeneousAnswers, instructions, code
Materials
knowledge
Chemistry
understanding
Physics
calculation
MacroPhysics
knowledge
IFEvalMBPPMacro
Student (initial)32.6356.4424.3837.8132.2554.9920.6935.98
Teacher67.4496.9456.0073.4677.2578.7057.6971.21
OPD55.5091.7139.7762.3359.2950.9531.4147.22
EOPD55.0092.7536.9461.5659.5650.2530.8846.90
OPDVR53.9892.8339.6762.1659.3552.3131.8347.83
ExOPD54.5693.2140.0862.6259.8149.7230.9546.83
FiRe-OPD54.6792.1738.3161.7258.9652.8930.7447.53
DuoOPD (ours)54.6592.4643.3363.4861.7553.3032.6949.24

DuoOPD also leads the five baselines in mean Macro on two additional task mixtures. The science mixture combines factual knowledge, understanding, and numerical calculation; the heterogeneous mixture combines physics answers, instruction following, and executable code. It leads the baselines on physics calculation in the science mixture and on all three tasks in the heterogeneous mixture.

IFEval uses prompt-level strict accuracy: a response must satisfy every instruction constraint. MBPP evaluates generated code. Values follow Tables 1 and 2 of the paper.

Outcome comparison for Qwen3 and Llama: DuoOPD increases student successes over OPD in biology, chemistry, and physics, on questions the fixed teacher both solves and fails.
On the same test questions with fixed teacher responses, DuoOPD improves over OPD both where the teacher succeeds and where it fails, in every domain for both families. Trained bars average three runs.

BibTeX

@article{yu2026duoopd,
  title={DuoOPD: Learning from Joint Teacher-Student Outcomes
         for Multi-Task On-Policy Distillation},
  author={Yu, Ao and Gao, Weibo and Zhou, Heng and Yue, Linan
          and Li, Rui and Liu, Suyi and Yan, Yu
          and Zhang, Yizhong and Liu, Qi},
  journal={arXiv preprint arXiv:2609.33711},
  year={2026}
}