source: arxiv machine learning: reward transport: property control in flow matching via noise-space alignment

level: research

flow matching models generate data by learning a mapping from noise to data points. the pairing of noise and data, called coupling, is usually seen as a computational detail. this work shows that coupling can embed control directly into the model. by aligning noise vectors with molecular properties during training, the learned flow field becomes steerable. the method, called reward transport, uses optimal transport to match a scalar noise coordinate with a target reward. at inference, changing this coordinate shifts the generated distribution toward higher-reward molecules.

the approach requires no reward model, gradient guidance, or extra computation at inference. in the coupling-preserving limit, thresholding the noise coordinate recovers the cross-entropy method's truncated reward distribution. this provides a principled, continuously adjustable control knob. experiments on zinc-250k and guacamol show that sweeping the scalar coordinate smoothly varies molecular properties like drug-likeness and synthetic accessibility. the generated molecules remain valid and diverse across the reward range.

reward transport offers a simple way to build controllable generative models for molecular design. it avoids the complexity of reinforcement learning or classifier guidance. the method is general and could apply to other domains where flow matching is used. by treating coupling as an alignment interface, it opens new possibilities for property-conditioned generation without inference overhead.

why it matters: it enables controllable molecular generation without extra models or gradients, simplifying ai-driven drug discovery and materials design.


source: arxiv machine learning: reward transport: property control in flow matching via noise-space alignment