[ICLR 2026]QeRL enables RL for 32B LLMs on a single H100 GPU.
Agent Reinforcement Trainer: train multi-step agents for real-world tasks using GRPO. Give your agents on-the-job training. Reinforcement learning for Qwen3.6, GPT-OSS, Llama, and more!
Awesome Reasoning LLM Tutorial/Survey/Guide
AgentFlow: In-the-Flow Agentic System Optimization
A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.