Explorer
- .. (Parent Directory)
- pdfs/
- a-formula-driven-survey-and-research-agenda-for-on-policy-distillation-2606.22793.md
- a-quantitative-approximation-framework-for-flow-distillation-in-diffus-2606.03820.md
- aliyunconsoleagent-training-web-agents-in-real-world-cloud-environment-2606.09447.md
- arkd-adaptive-reinforcement-learning-guided-bidirectional-kl-divergenc-2606.29869.md
- asyncopd-how-stale-can-on-policy-distillation-be-2606.24143.md
- atod-annealed-turn-aware-on-policy-distillation-for-multi-turn-autonom-2606.27814.md
- aura-internalizing-audio-understanding-into-llms-as-lora-2606.11033.md
- bahsd-bridging-the-long-tail-gap-via-adaptive-distillation-in-black-bo-2606.03091.md
- be-my-tutor-on-policy-co-distillation-for-mutual-llm-improvement-via-p-2606.14368.md
- behavior-cloning-is-not-all-you-need-the-optimality-of-on-policy-disti-2606.30923.md
- beyond-dark-knowledge-mixup-based-distillation-for-reliable-prediction-2606.12171.md
- beyond-output-matching-preserving-internal-geometry-in-nvfp4-llm-disti-2606.05682.md
- beyond-trajectory-imitation-strategy-guided-policy-optimization-for-ll-2606.24064.md
- bitnet-text-embeddings-2606.25674.md
- blockwise-policy-drift-gating-for-on-policy-distillation-2606.24084.md
- breaking-the-tokenizer-barrier-on-policy-distillation-across-model-fam-2606.09456.md
- building-multi-task-agentic-llms-via-two-phase-distillation-2606.30044.md
- candle-character-level-arabic-noise-deduplication-using-lightweight-en-2606.24758.md
- cartridges-at-scale-training-modular-kv-caches-over-large-document-col-2606.04557.md
- characterize-then-distill-mechanistic-reasoning-in-large-output-spaces-2606.06840.md
- compress-distill-reasoning-trace-compression-for-efficient-knowledge-d-2606.05988.md
- context-aware-distillation-and-ablation-for-text2dsl-2606.22578.md
- cross-modal-knowledge-distillation-without-paired-data-theoretical-fou-2606.10504.md
- danceopd-on-policy-generative-field-distillation-2606.27377.md
- data-efficient-autoregressive-to-diffusion-language-models-via-on-poli-2606.06712.md
- dense-supervision-sparse-updates-on-the-sparsity-and-geometry-of-on-po-2606.13657.md
- distill-on-a-diet-efficient-knowledge-distillation-via-learnable-data-2606.25488.md
- distill-once-adapt-life-long-exploring-dataset-distillation-for-contin-2606.20196.md
- distilling-answer-set-programming-rules-from-llms-for-neurosymbolic-vi-2606.03269.md
- distilling-drifting-transformers-with-representation-autoencoders-2606.15553.md
- distilling-examples-into-task-instructions-enhanced-in-context-learnin-2606.15641.md
- distilling-safe-llm-systems-via-soft-prompts-for-on-device-settings-2606.09388.md
- doc-to-atom-learning-to-compile-and-compose-memory-atoms-2606.12400.md
- dopd-dual-on-policy-distillation-2606.30626.md
- drift-difficulty-routing-self-distillation-with-rhythm-gated-explorati-2606.30345.md
- dudi-dual-signal-distillation-with-cross-lingual-verbalizer-2606.04694.md
- duomem-towards-capable-on-device-memory-agents-via-dual-space-distilla-2606.29961.md
- efficient-financial-language-understanding-via-distillation-with-synth-2606.18875.md
- egtr-review-efficient-evidence-grounded-scientific-peer-review-generat-2606.06025.md
- escaping-the-kl-agreement-trap-in-on-policy-distillation-2606.09471.md
- escaping-the-self-confirmation-trap-an-execute-distill-verify-paradigm-2606.24428.md
- fada-accessible-fetal-ultrasound-interpretation-and-annotation-with-a-2606.11106.md
- fast-speech-foundation-model-distillation-using-interleaved-stacking-2606.11766.md
- filter-then-reweight-rethinking-optimization-granularity-in-on-policy-2606.02684.md
- finding-the-evidence-discovering-decision-supporting-tokens-for-on-pol-2606.22830.md
- force-efficient-vla-reinforcement-fine-tuning-via-value-calibrated-war-2606.26006.md
- from-clicks-to-intent-cross-platform-session-embeddings-with-llm-disti-2606.26277.md
- gr2-technical-report-2606.31984.md
- harnessbridge-learnable-bidirectional-controller-for-llm-agent-harness-2606.12882.md
- heterophily-aware-adaptive-knowledge-distillation-for-hypergraph-neura-2606.08978.md
- hilda-hierarchical-distillation-with-diffusion-for-advancing-self-supe-2606.20189.md
- hump-kd-a-hybrid-uncertainty-aware-multi-stage-progressive-knowledge-d-2606.14684.md
- hyperdflash-hyper-connection-aligned-block-speculative-decoding-with-g-2606.26744.md
- improving-the-efficiency-and-effectiveness-of-llm-knowledge-distillati-2606.04650.md
- influcoder-distilling-decoders-gradient-influence-rankings-into-an-enc-2606.13668.md
- invariant-gradient-alignment-for-robust-reasoning-distillation-2606.05025.md
- k-forcing-joint-next-k-token-decoding-via-push-forward-language-modeli-2606.10820.md
- keep-policy-gradient-in-charge-sibling-guided-credit-distillation-for-2606.12634.md
- knowledge-cascade-reverse-knowledge-distillation-on-nonparametric-mult-2606.25927.md
- knowledge-distillation-for-visual-autoregressive-models-2606.06078.md
- knowledge-distillation-from-large-reasoning-models-to-compact-student-2606.31048.md
- language-models-need-sleep-learning-to-self-modify-and-consolidate-mem-2606.03979.md
- large-language-model-teaches-visual-students-cross-modality-transfer-o-2606.27527.md
- latent-reasoning-with-normalizing-flows-2606.06447.md
- learning-from-own-solutions-self-conditioned-credit-assignment-for-rei-2606.18810.md
- learning-from-the-self-future-on-policy-self-distillation-for-dllms-2606.18195.md
- learning-from-your-own-mistakes-constructing-learnable-micro-reflectiv-2606.18844.md
- learning-to-reason-by-analogy-via-retrieval-augmented-reinforcement-fi-2606.13680.md
- learning-visual-spatial-planning-from-symbolic-state-via-modality-gap-2606.06076.md
- leveraging-audio-llms-to-filter-speech-to-speech-training-data-2606.13507.md
- localizing-credit-at-the-divergence-path-conditioned-self-distillation-2606.15576.md
- localnav-distilling-frontier-vlms-and-embodied-rl-for-on-device-object-2606.27871.md
- lori-low-rank-distillation-for-implicit-reasoning-2606.05315.md
- marginal-advantage-accumulation-for-memory-driven-agent-self-evolution-2606.20475.md
- masked-language-flow-models-2606.27617.md
- metis-bridging-text-and-code-memory-for-self-evolving-agents-2606.24151.md
- mmg2skill-can-agents-distill-in-the-wild-guides-into-self-evolving-ski-2606.01993.md
- modf-sir-a-multi-agent-omni-modal-distilled-framework-for-social-intel-2606.12018.md
- mopd-multi-teacher-on-policy-distillation-for-capability-integration-i-2606.30406.md
- nebulaexp-8b-an-empirical-post-training-pipeline-via-full-scale-ablati-2606.26671.md
- odyssim-building-foundation-models-for-human-behavior-simulation-2606.14199.md
- omniopsd-rationale-privileged-on-policy-self-distillation-for-affectiv-2606.15920.md
- on-policy-distillation-with-curriculum-turn-level-guidance-for-multi-t-2606.15912.md
- on-policy-self-distillation-with-sampled-demonstrations-reduces-output-2606.26091.md
- on-the-geometry-of-on-policy-distillation-2606.07082.md
- on-the-position-bias-of-on-policy-distillation-2606.22600.md
- opd-evolver-cultivating-holistic-agent-evolver-via-on-policy-distillat-2606.17628.md
- open-swe-traces-advancing-dual-mode-multilingual-distillation-for-soft-2606.16038.md
- opid-on-policy-skill-distillation-for-agentic-reinforcement-learning-2606.26790.md
- oprd-on-policy-representation-distillation-2606.06021.md
- optimizing-teacher-student-partitioning-for-scalable-knowledge-distill-2606.27797.md
- padd-path-aligned-decompression-distillation-for-non-router-teacher-to-2606.10369.md
- parabridge-bridging-paralinguistic-perception-and-dialogue-behavior-in-2606.10581.md
- paved-with-true-intents-intent-aware-training-improves-llm-safety-clas-2606.27210.md
- pbsd-privileged-bayesian-self-distillation-for-long-horizon-credit-ass-2606.09348.md
- pguda-pressure-guided-unsupervised-domain-adaptation-with-cross-modal-2606.31349.md
- physics-guided-policy-optimization-with-self-distillation-2606.03620.md
- pool-select-refine-for-allocation-aware-generative-dataset-distillatio-2606.01920.md
- poweropd-stabilizing-on-policy-distillation-with-bounded-power-transfo-2606.17199.md
- pride-privileged-information-enhanced-distillation-for-empathetic-dial-2606.23124.md
- pythagoras-prover-advancing-efficient-formal-proving-via-augmented-lea-2606.12594.md
- q0-primitives-for-hyper-epoch-pretraining-2606.03938.md
- quantifying-subliminal-behavioral-transfer-ratios-in-language-model-di-2606.11270.md
- qval-cheaply-evaluating-dense-supervision-signals-for-long-horizon-llm-2606.32034.md
- qwen-image-2-0-rl-technical-report-2606.27608.md
- rank-aware-hyperbolic-alignment-for-vision-language-dataset-distillati-2606.29464.md
- rcem-robust-conversational-search-embedder-in-distributional-shift-2606.01697.md
- recover-lora-for-aggressive-quantization-reclaiming-accuracy-in-2-bit-2606.04238.md
- reinforcement-learning-from-rich-feedback-with-distributional-dagger-2606.05152.md
- renio-reweighting-negative-trajectory-importance-for-llm-on-policy-dis-2606.23104.md
- resaware-cross-environment-website-fingerprinting-via-resource-privile-2606.17462.md
- rethinking-continual-experience-internalization-for-self-evolving-llm-2606.04703.md
- rethinking-dataset-distillation-for-classification-do-distilled-sets-o-2606.18209.md
- rethinking-reward-supervision-rubric-conditioned-self-distillation-2606.19327.md
- rlcsd-reinforcement-learning-with-contrastive-on-policy-self-distillat-2606.11709.md
- road-vla-robust-online-adaptation-via-self-distillation-for-vision-lan-2606.25800.md
- robust-trajectory-distillation-hybrid-reweighting-meets-teacher-inspir-2606.29837.md
- rt-vla-real-time-vision-language-action-models-via-knowledge-distillat-2606.14010.md
- rubric-guided-self-distillation-post-training-without-rubric-verifiers-2606.12507.md
- safesteer-localized-on-policy-distillation-for-efficient-safety-alignm-2606.02530.md
- sage-opd-selective-agent-guided-intervention-for-multi-turn-on-policy-2606.19659.md
- sara-unlocking-multilingual-knowledge-in-mixture-of-experts-via-semant-2606.25821.md
- scalable-behaviour-cloning-on-browser-using-via-skill-distillation-2606.32014.md
- scaling-laws-for-task-specific-llm-distillation-2606.24747.md
- scaling-the-horizon-not-the-parameters-reaching-trillion-parameter-per-2606.30616.md
- seeing-before-reasoning-decoupling-perception-and-reasoning-for-shortc-2606.19120.md
- self-distilled-policy-gradient-2606.04036.md
- self-evaluation-is-already-there-eliciting-latent-judge-calibration-in-2606.05122.md
- self-study-reconsidered-the-hidden-fragility-of-learning-from-self-gen-2606.32002.md
- sg-opd-sign-gated-on-policy-distillation-via-sign-consistency-gating-a-2606.09304.md
- shard-safe-and-helpful-alignment-via-self-reframing-distillation-2606.15517.md
- siri-self-internalizing-reinforcement-learning-with-intrinsic-skills-f-2606.02355.md
- skill-disco-distilling-and-compiling-agent-traces-into-reusable-proced-2606.26669.md
- skill-guided-continuation-distillation-for-gui-agents-2606.18890.md
- spotattention-plug-in-block-sparse-routing-for-pretrained-long-context-2606.22874.md
- stabilizing-on-policy-distillation-for-mllm-reasoning-with-global-norm-2606.09091.md
- streamkl-fast-and-memory-efficient-kl-divergence-for-boosting-attentio-2606.20005.md
- structured-testbench-generation-for-llm-driven-hdl-design-and-verifica-2606.12983.md
- talk-text-attributed-graph-dataset-distillation-via-coupling-language-2606.22975.md
- taylor-calibrate-principled-initialization-for-hybrid-linear-attention-2606.16429.md
- teaching-the-way-not-the-answer-privileged-tutoring-distillation-for-m-2606.07000.md
- ternary-mamba-grouped-quantization-aware-training-of-w1-58a16-state-sp-2606.18114.md
- the-quality-utility-paradox-why-high-reward-data-impairs-small-model-m-2606.16152.md
- the-role-of-feedback-alignment-in-self-distillation-2606.11173.md
- towards-direct-latent-space-synthesis-for-parallel-branches-in-llm-age-2606.14672.md
- trust-the-right-teacher-quality-aware-self-distillation-for-gui-ground-2606.18101.md
- understanding-knowledge-distillation-in-post-training-when-it-helps-an-2606.22942.md
- uniego-proxies-as-mediators-for-unified-egocentric-video-representatio-2606.20559.md
- unsupervised-continual-clustering-via-forward-backward-knowledge-disti-2606.07474.md
- unsupervised-skill-discovery-for-agentic-data-analysis-2606.06416.md
- usad-2-0-scaling-representation-distillation-for-universal-audio-under-2606.06444.md
- vibethinker-3b-exploring-the-frontier-of-verifiable-reasoning-in-small-2606.16140.md
- vicur-visual-cues-as-recoverable-privilege-for-multimodal-on-policy-di-2606.05718.md
- visual-opsd-cross-modal-on-policy-self-distillation-for-efficient-unif-2606.18974.md
- what-does-it-mean-to-break-a-distillation-defense-2606.25059.md
- when-should-the-teacher-move-temporal-coupling-and-stability-in-self-o-2606.03532.md
- why-are-dmd-students-lazy-understanding-the-copying-behavior-in-few-st-2606.02237.md
- windom-self-family-distillation-for-small-model-gui-grounding-2606.25964.md
- youzhi-towards-high-concurrency-financial-llms-via-adaptive-gqa-to-mla-2606.05868.md
- zone-of-proximal-policy-optimization-teacher-in-prompts-not-gradients-2606.18216.md
- zooclaw-fashionsiglip2-distilled-fine-tuning-for-robust-fashion-retrie-2606.27708.md