Back to List
Notice:This resource is provided by a third-party author. Please review the code with AI tools or manually before use to ensure security and compatibility.
Pythonlasgroup/SDPO

SDPO

Reinforcement Learning via Self-Distillation (SDPO)

58.1/100
1.1KForks: 125
View on GitHubHomepage →
Loading report...

Similar Projects

TTRL

55

[NeurIPS 2025] TTRL: Test-Time Reinforcement Learning

Python1.1K

PageIndex

91

📑 PageIndex: Document Index for Vectorless, Reasoning-based RAG

Python35.6K

AReaL

88

The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.

Python5.7K

EasyR1

81

EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL

Python5.2K
Back to List