grpo-rl-training

davila7

Aktualisiert 11 days ago

293 Ansichten

18,478

1,685

18,478

Auf GitHub ansehen

DesignPost-TrainingReinforcement LearningGRPOTRLRLHFReward ModelingReasoningDPOPPOStructured Output

Über

Diese Fähigkeit bietet fachkundige Anleitung zur Implementierung von GRPO (Group Relative Policy Optimization) Reinforcement Learning Fine-Tuning mit der TRL-Bibliothek. Sie ist für das Training von Modellen für Aufgaben konzipiert, die strukturierte Ausgaben, überprüfbare Schlussfolgerungen oder objektive Korrektheitsmetriken wie bei Programmier- oder Mathematikaufgaben erfordern. Zu den Hauptmerkmalen gehören produktionsreife Workflows für benutzerdefinierte Belohnungsfunktionen und die Durchsetzung spezifischer Ausgabeformate.

Schnellinstallation

Claude Code

GitHub Repository

davila7/claude-code-templates

Pfad: cli-tool/components/skills/ai-research/post-training-grpo-rl-training

anthropicanthropic-claudeclaudeclaude-code

Verwandte Skills

openrlhf-training

Design

OpenRLHF is a high-performance RLHF training framework for fine-tuning large language models (7B-70B+ parameters) using methods like PPO, DPO, and GRPO. It leverages Ray for distributed architecture and vLLM for accelerated inference, achieving speeds 2x faster than alternatives like DeepSpeedChat. Use this skill when you need efficient, distributed RLHF training with optimized GPU resource sharing and ZeRO-3 support.

Skill ansehen

fine-tuning-with-trl

Andere

This skill enables fine-tuning of LLMs using TRL's reinforcement learning methods including SFT, DPO, and PPO for RLHF and preference alignment. It's designed for aligning models with human feedback and works with HuggingFace Transformers. Use it when you need to implement RLHF, optimize with rewards, or train from human preferences.

Skill ansehen

gptq

Andere

GPTQ is a 4-bit post-training quantization technique for LLMs that enables 4x memory reduction and 3-4x faster inference with minimal accuracy loss. It's ideal for deploying large models on consumer GPUs and integrates with transformers and PEFT for QLoRA fine-tuning. Use it when you need to fit 70B+ parameter models on limited hardware while maintaining performance.

Skill ansehen

instructor

Testen

Instructor is a structured output library that extracts validated data from LLM responses using Pydantic schemas. It automatically retries failed extractions and provides type-safe JSON parsing with streaming support. Use it when you need reliable, validated data extraction from LLMs like OpenAI or Anthropic.

Skill ansehen