Explorer
- .. (Parent Directory)
- 2606.01697 - RCEM - Robust Conversational Search EMbedder in Distributional Shift.pdf
- 2606.01920 - Pool-Select-Refine for Allocation-Aware Generative Dataset Distillation.pdf
- 2606.01993 - MMG2Skill - Can Agents Distill In-the-Wild Guides into Self-Evolving Skills.pdf
- 2606.02237 - Why Are DMD Students Lazy Understanding the Copying Behavior in Few-Step Distillation.pdf
- 2606.02355 - SIRI - Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training.pdf
- 2606.02530 - SafeSteer - Localized On-Policy Distillation for Efficient Safety Alignment.pdf
- 2606.02684 - Filter, Then Reweight - Rethinking Optimization Granularity in On-Policy Distillation.pdf
- 2606.03091 - BAHSD - Bridging the Long-tail Gap via Adaptive Distillation in Black-box Sequential Recommendation.pdf
- 2606.03269 - Distilling Answer-Set Programming Rules from LLMs for Neurosymbolic Visual Question Answering.pdf
- 2606.03532 - When Should the Teacher Move Temporal Coupling and Stability in Self On-Policy Distillation.pdf
- 2606.03620 - Physics-Guided Policy Optimization with Self-Distillation.pdf
- 2606.03820 - A Quantitative Approximation Framework for Flow Distillation in Diffusion Models.pdf
- 2606.03938 - q0 - Primitives for Hyper-Epoch Pretraining.pdf
- 2606.03979 - Language Models Need Sleep - Learning to Self-Modify and Consolidate Memories.pdf
- 2606.04036 - Self-Distilled Policy Gradient.pdf
- 2606.04238 - Recover-LoRA for Aggressive Quantization - Reclaiming Accuracy in 2-Bit Language Models via Low-Rank Adaptation with Knowledge Distillation on Synthetic Data.pdf
- 2606.04557 - Cartridges at Scale - Training Modular KV Caches over Large Document Collections.pdf
- 2606.04650 - Improving the Efficiency and Effectiveness of LLM Knowledge Distillation for Conversational Search.pdf
- 2606.04694 - DuDi - Dual-Signal Distillation with Cross-Lingual Verbalizer.pdf
- 2606.04703 - Rethinking Continual Experience Internalization for Self-Evolving LLM Agents.pdf
- 2606.05025 - Invariant Gradient Alignment for Robust Reasoning Distillation.pdf
- 2606.05122 - Self-Evaluation Is Already There - Eliciting Latent Judge Calibration in Base LLMs with Minimal Data.pdf
- 2606.05152 - Reinforcement Learning from Rich Feedback with Distributional DAgger.pdf
- 2606.05315 - LoRi - Low-Rank Distillation for Implicit Reasoning.pdf
- 2606.05682 - Beyond Output Matching - Preserving Internal Geometry in NVFP4 LLM Distillation.pdf
- 2606.05718 - ViCuR - Visual Cues as Recoverable Privilege for Multimodal On-Policy Distillation.pdf
- 2606.05868 - YouZhi - Towards High-Concurrency Financial LLMs via Adaptive GQA-to-MLA Transition.pdf
- 2606.05988 - Compress-Distill - Reasoning Trace Compression for Efficient Knowledge Distillation.pdf
- 2606.06021 - OPRD - On-Policy Representation Distillation.pdf
- 2606.06025 - EGTR-Review - Efficient Evidence-Grounded Scientific Peer Review Generation via Multi-Agent Teacher Distillation.pdf
- 2606.06076 - Learning Visual Spatial Planning from Symbolic State via Modality-Gap-Aware Self-Distillation.pdf
- 2606.06078 - Knowledge Distillation for Visual Autoregressive Models.pdf
- 2606.06416 - Unsupervised Skill Discovery for Agentic Data Analysis.pdf
- 2606.06444 - USAD 2.0 - Scaling Representation Distillation for Universal Audio Understanding.pdf
- 2606.06447 - Latent Reasoning with Normalizing Flows.pdf
- 2606.06712 - Data-Efficient Autoregressive-to-Diffusion Language Models via On-Policy Distillation.pdf
- 2606.06840 - Characterize Then Distill - Mechanistic Reasoning in Large Output Spaces.pdf
- 2606.07000 - Teaching the Way, Not the Answer - Privileged Tutoring Distillation for Multimodal Policy Optimization.pdf
- 2606.07082 - On the Geometry of On-Policy Distillation.pdf
- 2606.07474 - Unsupervised Continual Clustering via Forward-Backward Knowledge Distillation.pdf
- 2606.08978 - Heterophily-Aware Adaptive Knowledge Distillation for Hypergraph Neural Networks.pdf
- 2606.09091 - Stabilizing On-Policy Distillation for MLLM Reasoning with Global Normalization.pdf
- 2606.09304 - SG-OPD - Sign-Gated On-Policy Distillation via Sign-Consistency Gating and Phased Teacher Sampling.pdf
- 2606.09348 - PBSD - Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment.pdf
- 2606.09388 - Distilling Safe LLM Systems via Soft Prompts for On Device Settings.pdf
- 2606.09447 - AliyunConsoleAgent - Training Web Agents in Real-World Cloud Environments via Distillation and Reinforcement Learning.pdf
- 2606.09456 - Breaking the Tokenizer Barrier - On-Policy Distillation across Model Families.pdf
- 2606.09471 - Escaping the KL Agreement Trap in On-Policy Distillation.pdf
- 2606.10369 - PADD - Path-Aligned Decompression Distillation for Non-Router Teacher to Guide MoE Student Learning.pdf
- 2606.10504 - Cross-Modal Knowledge Distillation without Paired Data - Theoretical Foundation and Algorithm.pdf
- 2606.10581 - ParaBridge - Bridging Paralinguistic Perception and Dialogue Behavior in Speech Language Models.pdf
- 2606.10820 - K-Forcing - Joint Next-K-Token Decoding via Push-Forward Language Modeling.pdf
- 2606.11033 - AuRA - Internalizing Audio Understanding into LLMs as LoRA.pdf
- 2606.11106 - FADA - Accessible fetal ultrasound interpretation and annotation with a selectively distilled unified vision-language model.pdf
- 2606.11173 - The Role of Feedback Alignment in Self-Distillation.pdf
- 2606.11270 - Quantifying Subliminal Behavioral Transfer Ratios in Language Model Distillation.pdf
- 2606.11709 - RLCSD - Reinforcement Learning with Contrastive On-Policy Self-Distillation.pdf
- 2606.11766 - Fast Speech Foundation Model Distillation Using Interleaved Stacking.pdf
- 2606.12018 - MODF-SIR - A Multi-agent Omni-modal Distilled Framework for Social Intelligence Reasoning.pdf
- 2606.12171 - Beyond Dark Knowledge - Mixup-Based Distillation for Reliable Predictions.pdf
- 2606.12400 - Doc-to-Atom - Learning to Compile and Compose Memory Atoms.pdf
- 2606.12507 - Rubric-Guided Self-Distillation - Post-Training Without Rubric Verifiers.pdf
- 2606.12594 - Pythagoras-Prover - Advancing Efficient Formal Proving via Augmented Lean Formalisation.pdf
- 2606.12634 - Keep Policy Gradient in Charge - Sibling-Guided Credit Distillation for Long-Horizon Tool-Use Agents.pdf
- 2606.12882 - HarnessBridge - Learnable Bidirectional Controller for LLM Agent Harness.pdf
- 2606.12983 - Structured Testbench Generation for LLM-Driven HDL Design and Verification-Oriented Data Curation.pdf
- 2606.13507 - Leveraging Audio-LLMs to Filter Speech-to-Speech Training Data.pdf
- 2606.13657 - Dense Supervision, Sparse Updates - On the Sparsity and Geometry of On-Policy Distillation.pdf
- 2606.13668 - Influcoder - Distilling Decoders' Gradient Influence Rankings into an Encoder for Data Attribution.pdf
- 2606.13680 - Learning to Reason by Analogy via Retrieval-Augmented Reinforcement Fine-Tuning.pdf
- 2606.14010 - RT-VLA - Real-Time Vision-Language-Action Models via Knowledge Distillation.pdf
- 2606.14199 - OdysSim - Building Foundation Models for Human Behavior Simulation.pdf
- 2606.14368 - Be My Tutor - On-Policy Co-Distillation for Mutual LLM Improvement via Peer Feedback.pdf
- 2606.14672 - Towards Direct Latent-Space Synthesis for Parallel Branches in LLM-Agent Workflows.pdf
- 2606.14684 - HumP-KD - A Hybrid Uncertainty-Aware Multi-Stage Progressive Knowledge Distillation Framework for Efficient Fire Classification.pdf
- 2606.15517 - SHARD - Safe and Helpful Alignment via Self-Reframing Distillation.pdf
- 2606.15553 - Distilling Drifting Transformers with Representation Autoencoders.pdf
- 2606.15576 - Localizing Credit at the Divergence - Path-Conditioned Self-Distillation for LLM Reasoning.pdf
- 2606.15641 - Distilling Examples into Task Instructions - Enhanced In-Context Learning for Real-World B2B Conversations.pdf
- 2606.15912 - On-Policy Distillation with Curriculum Turn-level Guidance for Multi-turn Agents.pdf
- 2606.15920 - OmniOPSD - Rationale-Privileged On-Policy Self-Distillation for Affective Computing.pdf
- 2606.16038 - Open-SWE-Traces - Advancing Dual-Mode Multilingual Distillation for Software Engineering Agents.pdf
- 2606.16140 - VibeThinker-3B - Exploring the Frontier of Verifiable Reasoning in Small Language Models.pdf
- 2606.16152 - The Quality-Utility Paradox - Why High-Reward Data Impairs Small Model Mathematical Reasoning.pdf
- 2606.16429 - Taylor-Calibrate - Principled Initialization for Hybrid Linear Attention Distillation.pdf
- 2606.17199 - PowerOPD - Stabilizing On-Policy Distillation with Bounded Power Transformation.pdf
- 2606.17462 - ResAware - Cross-Environment Website Fingerprinting via Resource-Privileged Distillation.pdf
- 2606.17628 - OPD-Evolver - Cultivating Holistic Agent Evolver via On-Policy Distillation.pdf
- 2606.18101 - Trust the Right Teacher - Quality-Aware Self-Distillation for GUI Grounding.pdf
- 2606.18114 - Ternary Mamba - Grouped Quantization-Aware Training of W1.58A16 State Space Models.pdf
- 2606.18195 - Learning from the Self-future - On-policy Self-distillation for dLLMs.pdf
- 2606.18209 - Rethinking Dataset Distillation for Classification - Do Distilled Sets Outperform Coresets.pdf
- 2606.18216 - Zone of Proximal Policy Optimization - Teacher in Prompts, Not Gradients.pdf
- 2606.18810 - Learning from Own Solutions - Self-Conditioned Credit Assignment for Reinforcement Learning with Verifiable Rewards.pdf
- 2606.18844 - Learning from Your Own Mistakes - Constructing Learnable Micro-Reflective Trajectories for Self-Distillation.pdf
- 2606.18875 - Efficient Financial Language Understanding via Distillation with Synthetic Data.pdf
- 2606.18890 - Skill-Guided Continuation Distillation for GUI Agents.pdf
- 2606.18974 - Visual-OPSD - Cross-Modal On-Policy Self-Distillation for Efficient Unified Multimodal Reasoning.pdf
- 2606.19120 - Seeing Before Reasoning - Decoupling Perception and Reasoning for Shortcut-Resilient Multimodal On-Policy Self-Distillation.pdf
- 2606.19327 - Rethinking Reward Supervision - Rubric-Conditioned Self-Distillation.pdf
- 2606.19659 - SAGE-OPD - Selective Agent-Guided Intervention for Multi-Turn On-Policy Distillation.pdf
- 2606.20005 - StreamKL - Fast and Memory-Efficient KL Divergence for Boosting Attention Distillation.pdf
- 2606.20189 - HilDA - Hierarchical Distillation with Diffusion for Advancing Self-Supervised LiDAR Pre-training.pdf
- 2606.20196 - Distill Once, Adapt Life-Long - Exploring Dataset Distillation for Continual Test-Time Adaptation.pdf
- 2606.20475 - Marginal Advantage Accumulation for Memory-Driven Agent Self-Evolution.pdf
- 2606.20559 - UNIEGO - Proxies as Mediators for Unified Egocentric Video Representation Learning.pdf
- 2606.22578 - Context-Aware Distillation and Ablation for Text2DSL.pdf
- 2606.22600 - On the Position Bias of On-Policy Distillation.pdf
- 2606.22793 - A Formula-Driven Survey and Research Agenda for On-Policy Distillation.pdf
- 2606.22830 - Finding the Evidence - Discovering Decision-Supporting Tokens for On-Policy Reasoning Distillation.pdf
- 2606.22874 - SpotAttention - Plug-In Block-Sparse Routing for Pretrained Long-Context Transformers.pdf
- 2606.22942 - Understanding Knowledge Distillation in Post-Training - When It Helps and When It Fails.pdf
- 2606.22975 - TaLK - Text-attributed Graph Dataset Distillation via Coupling Language Model with Graph-Aware Kernel.pdf
- 2606.23104 - ReNIO - Reweighting Negative Trajectory Importance for LLM On-Policy Distillation.pdf
- 2606.23124 - PRIDE - Privileged Information-enhanced Distillation for Empathetic Dialogue Generation.pdf
- 2606.24064 - Beyond Trajectory Imitation - Strategy-Guided Policy Optimization for LLM Reasoning.pdf
- 2606.24084 - Blockwise Policy-Drift Gating for On-Policy Distillation.pdf
- 2606.24143 - AsyncOPD - How Stale Can On-Policy Distillation Be.pdf
- 2606.24151 - Metis - Bridging Text and Code Memory for Self-Evolving Agents.pdf
- 2606.24428 - Escaping the Self-Confirmation Trap - An Execute-Distill-Verify Paradigm for Agentic Experience Learning.pdf
- 2606.24747 - Scaling Laws for Task-Specific LLM Distillation.pdf
- 2606.24758 - CANDLE - Character-level Arabic Noise Deduplication using Lightweight Encoder.pdf
- 2606.25059 - What Does It Mean to Break a Distillation Defense.pdf
- 2606.25488 - Distill on a Diet - Efficient Knowledge Distillation via Learnable Data Pruning.pdf
- 2606.25674 - BitNet Text Embeddings.pdf
- 2606.25800 - ROAD-VLA - Robust Online Adaptation via Self-Distillation for Vision-Language-Action Models.pdf
- 2606.25821 - SARA - Unlocking Multilingual Knowledge in Mixture-of-Experts via Semantically Anchored Routing Alignment.pdf
- 2606.25927 - Knowledge Cascade - Reverse Knowledge Distillation on Nonparametric Multivariate Functional Estimation.pdf
- 2606.25964 - WinDOM - Self-Family Distillation for Small-Model GUI Grounding.pdf
- 2606.26006 - FORCE - Efficient VLA Reinforcement Fine-Tuning via Value-Calibrated Warm-up and Self-Distillation.pdf
- 2606.26091 - On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity.pdf
- 2606.26277 - From Clicks to Intent - Cross-Platform Session Embeddings with LLM-Distilled Taxonomy for Financial Services Recommendations.pdf
- 2606.26669 - SKILL-DISCO - Distilling and Compiling Agent Traces into Reusable Procedural Skills.pdf
- 2606.26671 - NebulaExp-8B - An Empirical Post-Training Pipeline via Full-Scale Ablation Research.pdf
- 2606.26744 - HyperDFlash - Hyper-Connection-Aligned Block Speculative Decoding with Gated Residual Reduction.pdf
- 2606.26790 - OPID - On-Policy Skill Distillation for Agentic Reinforcement Learning.pdf
- 2606.27210 - Paved with True Intents - Intent-Aware Training Improves LLM Safety Classification Across Training Regimes.pdf
- 2606.27377 - DanceOPD - On-Policy Generative Field Distillation.pdf
- 2606.27527 - Large Language Model Teaches Visual Students - Cross-Modality Transfer of Fine-Grained Conceptual Knowledge.pdf
- 2606.27608 - Qwen-Image-2.0-RL Technical Report.pdf
- 2606.27617 - Masked Language Flow Models.pdf
- 2606.27708 - ZooClaw-FashionSigLIP2 - Distilled Fine-tuning for Robust Fashion Retrieval.pdf
- 2606.27797 - Optimizing Teacher-Student Partitioning for Scalable Knowledge Distillation on HPC Systems.pdf
- 2606.27814 - ATOD - Annealed Turn-aware On-policy Distillation for Multi-turn Autonomous Agents.pdf
- 2606.27871 - LocalNav - Distilling Frontier VLMs and Embodied RL for On-Device Object Goal Navigation.pdf
- 2606.29464 - Rank-Aware Hyperbolic Alignment for Vision-Language Dataset Distillation.pdf
- 2606.29837 - Robust Trajectory Distillation - Hybrid Reweighting Meets Teacher-Inspired Targets.pdf
- 2606.29869 - ARKD - Adaptive Reinforcement Learning-Guided Bidirectional KL Divergence Distillation for Text Generation.pdf
- 2606.29961 - DuoMem - Towards Capable On-Device Memory Agents via Dual-Space Distillation.pdf
- 2606.30044 - Building Multi-Task Agentic LLMs via Two-Phase Distillation.pdf
- 2606.30345 - DRIFT - Difficulty Routing Self-DIstillation with Rhythm-Gated Exploration and Success BuFfer Training.pdf
- 2606.30406 - MOPD - Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training.pdf
- 2606.30616 - Scaling the Horizon, Not the Parameters - Reaching Trillion-Parameter Performance with a 35B Agent.pdf
- 2606.30626 - DOPD - Dual On-policy Distillation.pdf
- 2606.30923 - Behavior Cloning is Not All You Need - The Optimality of On-Policy Distillation for Noisy Expert Feedback.pdf
- 2606.31048 - Knowledge Distillation from Large Reasoning Models to Compact Student Models - A Case Study on the John O Bryan Mathematics Competition.pdf
- 2606.31349.pdf
- 2606.31984 - GR2 Technical Report.pdf
- 2606.32002 - Self-Study Reconsidered - The Hidden Fragility of Learning from Self-Generated QA.pdf
- 2606.32014 - Scalable Behaviour Cloning on Browser Using via Skill Distillation.pdf
- 2606.32034 - QVal - Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents.pdf