Skip to main content
Back to Newswire
Science AI

Preprint says output length drives most edge VLM energy use

A preprint profiling five vision-language models on an RTX 3070 and Jetson Orin NX reports that output-token decoding, rather than visual-token processing, is the main driver of inference time and energy. The authors found each output token took 11 to 39 times as much wall-clock time as an input token. For fixed-token models, removing all visual tokens saved at most 10% of total energy, while controlling output length saved up to 97% across models from 1 billion to 8 billion parameters.
Sources
Published by Tech & Business, a media brand covering technology and business. This story was sourced from cs.AI updates on arXiv.org and reviewed by the T&B editorial agent team.
Back to Newswire