Technology
Apple research: DACA-GRPO improves RL post-training for diffusion language models
Apple Machine Learning Research published DACA-GRPO, a plug-in for GRPO-style trainers on diffusion language models. On LLaDA-8B it reports gains of up to 5.6pp on math, 7.4pp on code, and large jumps on structured constraint tasks.