[research]By ByteBulletin Editor
RL or SFT? New Research Teases Apart What Actually Drives Reasoning in LLMs
A new arxiv study compares reinforcement learning against supervised fine-tuning to isolate which training method truly boosts reasoning performance in large language models.
