News
AI Summary
18 May 20262 Dhuʻl-Hijjah 1447 AH
Tsinghua and Alibaba Joint Paper Introduces ViT³: A Vision Transformer with Linear Complexity — CVPR 2026 Oral

Tsinghua and Alibaba Joint Paper Introduces ViT³: A Vision Transformer with Linear Complexity — CVPR 2026 Oral

Tsinghua University and Alibaba have introduced ViT³, a novel vision transformer architecture achieving linear computational complexity. This advancement, presented at CVPR 2026, aims to make high-resolution image understanding feasible on edge devices. The model reinterprets the attention mechanism through Test-Time Training, reducing the traditional quadratic complexity associated with increasing image resolution. By constructing a lightweight internal model during inference, ViT³ enables more efficient information processing.

Follow these topics

Sign in to follow the topics that matter to you

Sign in to follow

This summary is generated with AI and receives periodic editorial review. Refer to the original source for full details.

0
0 reading now

Insight Score

Rate to unlock

Sign in to react, rate, and save. Sign In