NOW LET US – AI RAG SaaS Studio TP.HCM
NOW LET US
Digital Product Studio
Back to news
AGENTIC-SYSTEMS...1 min read

Fusion is not one-size-fits-all: Cross-Modal Representation Alignment for Time-to-Event Modeling

Share
NOW LET US Article – Fusion is not one-size-fits-all: Cross-Modal Representation Alignment for Time-to-Event Modeling

Researchers introduce a foundation model-driven framework for cross-modal representation alignment between CT imaging and longitudinal EHR data to improve time-to-event prediction. The study demonstrates that task-aware multimodal alignment is essential for robust generalization and scalable clinical deployment.

Computer Science > Artificial Intelligence

Title:Fusion is not one-size-fits-all: Cross-Modal Representation Alignment for Time-to-Event Modeling

View PDF HTML (experimental)Abstract:Accurate time-to-event (TTE) prediction from multimodal clinical data remains challenging due to modality imbalance and distribution shift. We introduce a foundation model-driven framework for cross-modal representation alignment between CT imaging and longitudinal EHR data, designed to generalize across tasks and institutions. CT and EHR modalities are encoded independently using domain-specific foundation models and aligned in a shared latent space through four principled fusion strategies: late fusion, contrastive alignment, cross-attention, and co-attention. We evaluate two clinically distinct TTE tasks: pulmonary embolism (PE) mortality and cardiovascular disease (CVD) outcomes, on large-scale multi-institutional cohorts (PE: N=3,099 train; 1,098 internal; 435 external; CVD: N=2,951 train; 837 internal; 682 external). Fusion consistently improves concordance index by 1.5-5.4% over unimodal baselines when modalities contribute comparably. Overall, contrastive multimodal fusion, particularly with CLMBR representations, provided the most consistent and statistically robust improvements, especially for PE mortality prediction. For MACE, cross-attention (one-hot) achieved the highest internal performance and image-guided co-attention achieved the best external performance. We therefore introduce a generalizable foundation model-based cross-modal alignment framework and provide the first systematic analysis of fusion behavior under modality imbalance in TTE prediction. Our results establish task-aware multimodal alignment as a necessary design principle for robust generalization and scalable clinical deployment.

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

© 2026 Now Let Us. All rights reserved.

Source: arXiv cs.AI Recent

Advertisement
Ad slot ready: 5887729102

More in this category

NOW LET US Related – Metric Match: A Subset Selection Approach to Evaluating LLM Judge Reliability

agentic-systems

Metric Match: A Subset Selection Approach to Evaluating LLM Judge Reliability

Researchers have developed Metric Match, a novel method to estimate the reliability of LLM judges using limited human annotations. By selecting an optimal subset of samples, it reduces annotation needs by 32.5% and significantly cuts down evaluation costs.

NOW LET US Related – Semantics-Enhanced Retrieval-Augmented Time Series Forecasting

agentic-systems

Semantics-Enhanced Retrieval-Augmented Time Series Forecasting

Researchers have introduced SERAF, a novel multimodal time series forecasting framework that addresses the limitations of traditional methods under non-stationarity by leveraging dual retrieval over both numerical data and self-generated textual descriptions.

NOW LET US Related – AI Engram: In Search of Memory Traces in Artificial Intelligence

agentic-systems

AI Engram: In Search of Memory Traces in Artificial Intelligence

Researchers introduce 'AI Engram', a geometric framework to identify and isolate individual memory traces in deep neural networks. This biologically-inspired approach enables surgical manipulation of learned knowledge through simple linear arithmetic without iterative optimization.

NOW LET US Related – Risk-Aware LLM Agents for Geospatial Data Retrieval: Design and Preliminary Adversarial Evaluation

agentic-systems

Risk-Aware LLM Agents for Geospatial Data Retrieval: Design and Preliminary Adversarial Evaluation

Researchers have presented an LLM-driven framework that simplifies the retrieval of remote sensing data from cloud-based geospatial catalogs using natural language. The system integrates three specialized AI agents to optimize performance and mitigate adversarial API manipulation risks.

NOW LET US Related – Attribute Inference from Interactive Targeted Ads

agentic-systems

Attribute Inference from Interactive Targeted Ads

A new study models how interactive targeted advertising can act as a channel for attribute inference, allowing advertisers to deduce sensitive user data. The researchers propose defense mechanisms like aggregate reporting and randomized disclosure to mitigate these privacy risks.

NOW LET US Related – Visual-Seeker: Towards Visual-Native Multimodal Agentic Search via Active Visual Reasoning

agentic-systems

Visual-Seeker: Towards Visual-Native Multimodal Agentic Search via Active Visual Reasoning

Researchers have introduced Visual-Seeker, a pioneering visual-native multimodal search agent that addresses the factual grounding limitations of current AI models. By leveraging active visual reasoning, it outperforms several proprietary models in complex web search tasks.

EXPLORE TOPICS

Discover All Categories

Deep dive into the specific technology sectors that matter most to you.