Reinforcement Learning with Human Feedback (RLHF)
TRL
RLHF on GPT-2
สอนให้ Model Generate ข้อความเชิงบวก (Positive Sentiment) ได้มากขึ้นด้วย PPO
Colab Notebook — RLHF Positive Sentiment ด้วย PPO บน GPT-2colab.research.google.com
สอนให้ Model Generate ข้อความในเชิงบวก กลางๆ หรือเชิงลบ (Controlled Sentiment) โดยการกำหนด Prefix ใน Input
Colab Notebook — RLHF Controlled Sentiment บน GPT-2colab.research.google.com