OmniSIFT: Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language Models
IntermediateYue Ding, Yiyan Ji et al.Feb 4arXiv
OmniSIFT is a new way to shrink (compress) audio and video tokens so omni-modal language models can think faster without forgetting important details.
#Omni-LLM#token compression#modality-asymmetric