Archives AI News

Rethinking Diffusion Model in High Dimension

arXiv:2503.08643v3 Announce Type: replace-cross Abstract: Curse of Dimensionality is an unavoidable challenge in statistical probability models, yet diffusion models seem to overcome this limitation, achieving impressive results in high-dimensional data generation. Diffusion models assume that they can learn the statistical…

TDHook: A Lightweight Framework for Interpretability

arXiv:2509.25475v1 Announce Type: new Abstract: Interpretability of Deep Neural Networks (DNNs) is a growing field driven by the study of vision and language models. Yet, some use cases, like image captioning, or domains like Deep Reinforcement Learning (DRL), require complex…

Unified Cross-Modal Image Synthesis with Hierarchical Mixture of Product-of-Experts

arXiv:2410.19378v2 Announce Type: replace-cross Abstract: We propose a deep mixture of multimodal hierarchical variational auto-encoders called MMHVAE that synthesizes missing images from observed images in different modalities. MMHVAE’s design focuses on tackling four challenges: (i) creating a complex latent representation…

Discovering and Steering Interpretable Concepts in Large Generative Music Models

arXiv:2505.18186v2 Announce Type: replace-cross Abstract: The fidelity with which neural networks can now generate content such as music presents a scientific opportunity: these systems appear to have learned implicit theories of such content’s structure through statistical learning alone. This offers…

FW-Merging: Scaling Model Merging with Frank-Wolfe Optimization

arXiv:2503.12649v3 Announce Type: replace Abstract: Model merging has emerged as a promising approach for multi-task learning (MTL), offering a data-efficient alternative to conventional fine-tuning. However, with the rapid development of the open-source AI ecosystem and the increasing availability of fine-tuned…

Ban&Pick: Ehancing Performance and Efficiency of MoE-LLMs via Smarter Routing

arXiv:2509.06346v2 Announce Type: replace Abstract: Sparse Mixture-of-Experts (MoE) has become a key architecture for scaling large language models (LLMs) efficiently. Recent fine-grained MoE designs introduce hundreds of experts per layer, with multiple experts activated per token, enabling stronger specialization. However,…

Are neural scaling laws leading quantum chemistry astray?

arXiv:2509.26397v1 Announce Type: cross Abstract: Neural scaling laws are driving the machine learning community toward training ever-larger foundation models across domains, assuring high accuracy and transferable representations for extrapolative tasks. We test this promise in quantum chemistry by scaling model…