[Paper] GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention

Published: 6 days ago (June 4, 2026 at 10:52 AM EDT)

2 min read

Source: arXiv

Source: arXiv - 2606.06249v1

Overview

Transformer-based multimodal models rely on attention mechanisms to integrate information across heterogeneous modalities. Despite their success, existing multimodal attention formulations compute their scores through collections of pairwise dot-product interactions or by concatenating all the modalities into the keys, even when multiple modalities should be jointly involved. As a consequence, current approaches either incur quadratic complexity in the number of modalities or fail to explicitly model interactions that depend on the joint configuration of multiple representations. In this work, we introduce the Volumetric Multimodal cross-Attention (VMA), a novel cross-attention mechanism in which attention scores are defined as a function of the joint geometry of a query and multiple modality-specific keys. VMA computes the volume spanned by query and key vectors across multiple modalities, capturing joint multimodal dependencies beyond pairwise similarity, enabling native modeling of any-order modality interactions. We integrate VMA into our novel multimodal transformer architecture, named GRAMformer, explicitly designed to integrate any number of modalities. We evaluate the proposed model on multimodal learning tasks, demonstrating improved effectiveness and efficiency.

Key Contributions

This paper presents research in the following areas:

cs.CV
cs.LG

Methodology

Please refer to the full paper for detailed methodology.

Practical Implications

This research contributes to the advancement of cs.CV.

Authors

Giordano Cicchetti
Eleonora Grassucci
Danilo Comminiello

Paper Information

arXiv ID: 2606.06249v1
Categories: cs.CV, cs.LG
Published: June 4, 2026
PDF: Download PDF

[Paper] GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention

Overview

Key Contributions

Methodology

Practical Implications

Authors

Paper Information

Related posts

[Paper] MemDreamer: Decoupling Perception and Reasoning for Long Video Understanding via Hierarchical Graph Memory and Agentic Retrieval Mechanism

[Paper] Planning-aligned Token Compression for Long-Context Autonomous Driving

[Paper] TEVI: Text-Conditioned Editing of Visual Representations via Sparse Autoencoders for Improved Vision-Language Alignment

[Paper] Watch, Remember, Reason: Human-View Video Understanding with MLLMs