Blog

Notes on post-training, RLHF, evaluation, and human-centered AI.

DPO or PPO? Choose by Data, Not Fashion

A practical framework for choosing between offline preference optimization and online reinforcement learning based on data freshness, feedback quality, infrastructure, and failure modes.

  • Post-Training
  • RLHF
  • DPO
  • PPO
Read essay

Why Online Learning Still Matters in Post-Training

Offline preference fitting can teach a model from yesterday’s comparisons. Online learning becomes valuable when the model must learn from the distribution its own improving policy creates.

  • RLHF
  • Reinforcement Learning
  • Post-Training
Read essay