Daily Paper Cast
Vision Bridge Transformer at Scale
02 December 2025 21:48 Jingwen Liang, Gengyu Wang
Listen to episode
About this episode
🤗 Upvotes: 31 | cs.CV, cs.AI
<strong>Authors:</strong><br />
Zhenxiong Tan, Zeqing Wang, Xingyi Yang, Songhua Liu, Xinchao Wang</p>
<strong>Title:</strong><br />
Vision Bridge Transformer at Scale</p>
<strong>Arxiv:</strong><br />
<a href="http://arxiv.org/abs/2511.23199v1">http://arxiv.org/abs/2511.23199v1</a></p>
<strong>Abstract:</strong><br />
We introduce Vision Bridge Transformer (ViBT), a large-scale instantiation of Brownian Bridge Models designed for conditional generation. Unlike traditional diffusion models that transform noise into data, Bridge Models directly model the trajectory between inputs and outputs, creating an efficient data-to-data translation paradigm. By scaling these models to 20B and 1.3B parameters, we demonstrate their effectiveness for image and video translation tasks. To support this scale, we adopt a Transformer architecture and propose a variance-stabilized velocity-matching objective for robust training. Together, these advances highlight the power of scaling Bridge Models for instruction-based image editing and complex video translation
More AI podcast episodes
Browse all →Want to find AI jobs?
Join thousands of AI professionals finding their next opportunity