OpenCoF: Learning to Reason Through Video Generation
Paper • 2607.08763 • Published • 26
Wan-CoF is the data-only model from OpenCoF: Learning to Reason Through Video Generation. It is obtained by LoRA fine-tuning Wan2.2-I2V-A14B on OpenCoF-17K, without adding a reasoning-specific architecture.
Both adapters are required because the base model has separate denoising experts:
high_noise_lora/model.safetensors
low_noise_lora/model.safetensors
The adapters do not include the Wan2.2-I2V-A14B base weights.
Headline results reported in the paper:
| Benchmark | Metric | Wan2.2-I2V-A14B | Wan-CoF |
|---|---|---|---|
| MME-CoF | Overall (0–4) | 1.00 | 1.30 |
| VIPER | POC@1.0 | 3.3 | 7.5 |
| RULER-Bench | Overall, I2V subset (0–100) | 55.8 | 56.8 |
See the paper for the complete protocol and per-category results.
The adapters are released under Apache-2.0 and require the separately licensed Wan2.2-I2V-A14B base model. Please cite both OpenCoF and Wan when using them.
@article{chen2026opencof,
title = {OpenCoF: Learning to Reason Through Video Generation},
author = {Chen, Xinyan and Guo, Ziyu and Zhang, Renrui and Jiang, Dongzhi and Li, Hongsheng},
journal = {arXiv preprint arXiv:2607.08763},
year = {2026}
}
Base model
Wan-AI/Wan2.2-I2V-A14B