ACACIA Associate-Team

ACcelerAting throughput and reduCIng resource usage for generative AI


ACACIA is an Inria associate-team for the period 2026–2028 between:

The team is also supported by the International Laboratory on Learning Systems.

Organization

Members

  • Oana Balmau, DISCS Lab, McGill University, Montreal (PI-McGill)
  • Olivier Beaumont, Inria Topal team, Bordeaux
  • Mark Coates, Networks Research Lab, McGill University, Montreal
  • Lionel Eyraud-Dubois-Dubois, Inria Topal team, Bordeaux
  • Thomas Herault, Inria Topal team, Bordeaux
  • Laercio Lima Pilla, Inria Topal team, Bordeaux
  • Loris Marchal, Inria ROMA Team, Lyon (PI-Inria)

Visits and Meetings

The virtual kick-off meeting of the associate team was held on July 27, 2026. First visits are planned for Fall 2026.

Scientific context and objectives

Generative AI has rapidly become unavoidable, with widely used applications such as code assistants, text analysis, information retrieval, and translation. The core architecture of generative AI, relying on the transformer framework, inherently requires sequential token generation, leading to high computing times. Furthermore, the increasing complexity of models, requires a very large memory capacity in order to store them. Recent usages such as “reasoning” models even further push the memory demand because of their very large context requirements. When model weights and context exceed the available memory, frequent I/O operations become necessary, degrading performance and efficiency. In the end, a substantial and growing energy consumption is associated with the use of these models, raising concerns about both environmental impact and operational feasibility.

Most recent improvements in LLM quality have been mainly driven by scaling up the model. While it has led to remarkable results, this calls for even larger computing power and memory size. To ensure that generative AI remains accessible, efficient, and environmentally responsible, there is an urgent need to optimize resource usage. This can be achieved through architectural innovations that reduce the computational load, smarter memory management, and more efficient data loading strategies.

Our proposal builds upon existing techniques from the literature designed to accelerate the inference and/or training of large language models. We aim to adapt and enhance three of these approaches for our specific objectives: the use of multiple models of varying cost and quality, organized in “cascades”; the application of low-rank adapters to specialize models for particular tasks; and the adoption of the mixture-of-experts paradigm to reduce the memory requirement during inference.

More information to come.

Reference publications from the associate-team members

  • Antonios Valkanas, Soumyasundar Pal, Pavel Rumiantsev, Yingxue Zhang, and Mark Coates. C3PO: Optimized large language model cascades with probabilistic cost constraints for reasoning. In NeurIPS, 2025.
  • Florence Regol, Joseph Cotnareanu, Theodore Glavas, and Mark Coates. Is the acquisition worth the cost? Surrogate losses for consistent two-stage classifiers. In Poster at NeurIPS, 2025.
  • Florence regol, Joud Chataoui, and Mark Coates. Jointly-learned exit and inference for a dynamic neural network. In The Twelfth International Conference on Learning Representations (ICLR), 2024.
  • Spyros Angelopoulos, Loris Marchal, Adrien Obrecht, and Bertrand Simon. Cache management for mixture-of-experts llms. In Euro-Par 2025: Parallel Processing, volume 15902 of Lecture Notes in Computer Science, pages 18–32. Springer, 2025.
  • Olivier Beaumont, Raphaël Bourgouin, Maxime Darrin, Loris Marchal, and Pablo Piantanida. Leveraging expert usage to speed up LLM inference with expert parallelism. In Euro-Par 2025: Parallel Processing, volume 15900 of Lecture Notes in Computer Science, pages 145–158. Springer, 2025.
  • Zachary Doucet, Rishi Sharma, Martijn de Vos, Rafael Pires, Anne-Marie Kermarrec, and Oana Balmau. Harmoeny: Efficient multi-GPU inference of MoE models, Arxiv Preprint, 2025.
  • Oana Balmau, Anne-Marie Kermarrec, Rafael Pires, André Loureiro Espı́rito Santo, Martijn de Vos, and Milos Vujasinovic. Accelerating MoE model inference with expert sharding. In Eiko Yoneki and Amir H. Payberah, editors, Proceedings of the 5th Workshop on Machine Learning and Systems (EuroMLSys), pages 192–199. ACM, 2025.
  • Rahma Nouaji, Stella Bitchebe, Ricardo Macedo, and Oana Balmau. MinatoLoader: Accelerating machine learning training through efficient data preprocessing, Arxiv Preprint, 2025.
  • Dufy Teguia, Jiaxuan Chen, Stella Bitchebe, Oana Balmau, and Alain Tchana. vpim: Processing-in-memory virtualization. In Proceedings of the 25th International Middleware Conference, Middleware '24, page 417–430, 2024.
  • Adrien Aguila-Multner, Olivier Beaumont, Lionel Eyraud-Dubois, and Julia Gusak. Optimized Forward-Backward Rematerialization for Memory-Efficient Pipeline Parallel Training. preprint, July 2025.
  • Olivier Beaumont, Lionel Eyraud-Dubois, Julien Herrmann, Alexis Joly, and Alena Shilova. Optimal re-materialization strategies for heterogeneous chains: How to train deep neural networks with limited memory. ACM Transactions on Mathematical Software, 50(2):1–38, 2024.
  • Julia Gusak, Xunyi Zhao, Théotime Le Hellard, Zhe Li, Lionel Eyraud-Dubois, and Olivier Beaumont. Hiremate: Hierarchical approach for efficient re-materialization of large neural networks. In International Conference on Machine Learning (ICML), 2025.
  • Xunyi Zhao, Théotime Le Hellard, Lionel Eyraud-Dubois, Julia Gusak, and Olivier Beaumont. Rockmate: an efficient, fast, automatic and generic tool for re-materialization in pytorch. In International Conference on Machine Learning (ICML), 2023.

Author: Loris Marchal

Created: 2026-07-30 Thu 11:42

Emacs 25.3.50.1 (Org mode 8.2.10)

Validate