Saltar al contenido principal
Volver al leaderboard

dev

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning

📄 Paper Trending (73⬆️): Large language models are increasingly trained as interactive agents for long-horizon tasks involving multi-turn interaction, tool use, and environment feedback. Outcome-based reinforcement learning (RL) provides a practical optimization paradigm, but its sparse trajectory-level rewards offer limite...

Fuente: HuggingFace_Papers

Qué hace

📄 Paper Trending (73⬆️): Large language models are increasingly trained as interactive agents for long-horizon tasks involving multi-turn interaction, tool use, and environment feedback. Outcome-based reinforcement learning (RL) provides a practical optimization paradigm, but its sparse trajectory-level rewards offer limite...

Para quién es

No disponible

Precio

No disponible

Puntos fuertes

  • No disponible

Límites

  • No disponible

¿Quieres integrarlo en tu negocio? Reserva una llamada gratuita

Reservar