Life sciences · Preprint
arXiv · September 9, 2026
Posted before peer review. The findings may change or fail to hold.
This is a preprint describing FlowCPO, an offline preference optimization method for flow models that uses a forward-KL objective with contrastive flow matching. The method shows higher performance than FlowDPO on in-domain image generation tasks (GenEval 0.84, OCR 0.87) but mixed results out-of-domain, and has not been peer reviewed.
Preprint. Intervention: FlowCPO offline forward-KL objective using both preferred and dispreferred samples with contrastive flow matching. Compared with: FlowDPO and RFT (reward fine-tuning) baselines.
FlowCPO achieves GenEval score of 0.84 and OCR score of 0.87 in-domain at CFG 3.0, versus 0.81 and 0.74 for FlowDPO Forward-KL objective is bounded by contrastive flow matching loss under explicit regularity conditions for linear interpolation Signed regression loss of simplified FlowDPO can be unbounded below, whereas proposed loss is nonnegative
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
This is an unrefereed arXiv preprint introducing a machine learning method for preference alignment in generative models, with theoretical analysis and empirical results that have not undergone peer review.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Preference alignment for flow and diffusion models now spans online reinforcement learning and offline preference optimization, but the relation between these methods remains unclear. In particular, existing forward-process alignment methods require fresh samples from the current model, while offline methods based on fixed preference pairs rely primarily on positive-only fine-tuning or DPO-style likelihood-ratio surrogates. We organize these approaches through a divergence-based framework and introduce FlowCPO, an offline forward-KL objective that uses both preferred and dispreferred samples without online rollouts. For linear interpolation, we show under explicit regularity conditions that the forward-KL objective is bounded by a contrastive flow matching loss, yielding a tractable surrogate on fixed data. We further show that this loss is nonnegative, whereas the signed regression loss of simplified FlowDPO can be unbounded below. In the in-domain setting, FlowCPO achieves higher mean GenEval and OCR scores than the evaluated baselines, reaching 0.84 and 0.87 versus 0.81 and 0.74 for FlowDPO at CFG 3.0. In the out-of-domain setting, the results are mixed, with the best GenEval result but lower reward scores than RFT on several metrics.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.