Skip to main navigation Skip to search Skip to main content

HiRED: attention-guided token dropping for efficient inference of high-resolution vision-language models

  • Kazi Hasan Ibn Arif
  • , JinYi Joon
  • , Dimitrios S. Nikolopoulos
  • , Hans Vandierendonck
  • , Deepu John
  • , Bo Ji

Research output: Chapter in Book/Report/Conference proceedingConference contribution

Abstract

High-resolution Vision-Language Models (VLMs) are widely used in multimodal tasks to enhance accuracy by preserving detailed image information. However, these models often generate an excessive number of visual tokens due to the need to encode multiple partitions of a high-resolution image input. Processing such a large number of visual tokens poses significant computational challenges, particularly for resource-constrained commodity GPUs. To address this challenge, we propose High-Resolution Early Dropping (HiRED), a plug-and-play token-dropping method designed to operate within a fixed token budget. HiRED leverages the attention of CLS token in the vision transformer (ViT) to assess the visual content of the image partitions and allocate an optimal token budget for each partition accordingly. The most informative visual tokens from each partition within the allocated budget are then selected and passed to the subsequent Large Language Model (LLM). We showed that HiRED achieves superior accuracy and performance, compared to existing token-dropping methods. Empirically, HiRED-20% (i.e., a 20% token budget) on LLaVA-Next-7B achieves a 4.7x increase in token generation throughput, reduces response latency by 78%, and saves 14% of GPU memory for single inference on an NVIDIA TESLA P40 (24 GB). For larger batch sizes (e.g., 4), HiRED-20% prevents out-of-memory errors by cutting memory usage by 30%, while preserving throughput and latency benefits.

Original languageEnglish
Title of host publication39th AAAI Conference on Artificial Intelligence (AAAI-25): Proceedings
EditorsToby Walsh, Julie Shah, Zico Kolter
PublisherAAAI Press
Pages1773-1781
Number of pages9
ISBN (Print)9781577358978
DOIs
Publication statusPublished - 11 Apr 2025
Event39th Annual AAAI Conference on Artificial Intelligence 2025 - Philadelphia, United States
Duration: 25 Feb 202504 Mar 2025
Conference number: 39
https://aaai.org/conference/aaai/aaai-25/

Publication series

NameProceedings of the AAAI Conference on Artificial Intelligence
PublisherAAAI Press
Number2
Volume39
ISSN (Print)2159-5399
ISSN (Electronic)2374-3468

Conference

Conference39th Annual AAAI Conference on Artificial Intelligence 2025
Abbreviated titleAAAI-25
Country/TerritoryUnited States
CityPhiladelphia
Period25/02/202504/03/2025
Internet address

Keywords

  • token dropping
  • high-resolution
  • vision-language models

Fingerprint

Dive into the research topics of 'HiRED: attention-guided token dropping for efficient inference of high-resolution vision-language models'. Together they form a unique fingerprint.

Cite this