Meta paper uses RL to match how often models produce certain outputs

MaD-RL treats post-training as distribution matching, not only per-answer reward. Researchers listed on a Meta AI publication dated Sept. 24 introduced MaD-RL, a reinforcement-learning method meant to steer the mix of language-model outputs toward a target distribution. Not just a…





