arXiv:2603.15685v2 Announce Type: replace-cross Abstract: Omnimodal large language models (OmniLLMs) jointly process audio and visual streams, but the resulting long multimodal token sequences make inference prohibitively expensive. Existing compression methods typically rely on fixed window partitioning and attention-based pruning, which overlook the piecewise semantic structure…
Thank you for reading this post, don't forget to subscribe!
Source: cs.AI updates on arXiv.org
Automatically aggregated summary — full article and all rights belong to the original publisher.