NVIDIA FLARE Expands Federated Learning with Kubernetes, Slurm

Binance
Paxful




Jessie A Ellis
Sep 15, 2026 17:05

NVIDIA FLARE 2.9 adds Slurm support, enabling scalable federated learning across heterogeneous infrastructures like Docker, Kubernetes, and HPC clusters.



NVIDIA FLARE Expands Federated Learning with Kubernetes, Slurm

NVIDIA has unveiled an upgrade to its federated learning platform, FLARE, introducing support for Slurm scheduling and expanding compatibility across Docker and Kubernetes environments. This advancement allows organizations to scale federated learning (FL) projects across heterogeneous infrastructures without forcing standardization on participating sites. FLARE 2.9, released in September 2026, is poised to address one of FL’s persistent roadblocks: operational complexity in multi-organization collaboration.

Federated learning is a decentralized machine-learning approach where sensitive data remains local while models are trained collaboratively. Initially popularized by Google in 2016, FL is now widely adopted in privacy-sensitive industries like healthcare and finance. NVIDIA FLARE is a key tool in this ecosystem, enabling secure, scalable execution of FL workflows.

Solving Infrastructure Fragmentation

In many FL use cases, participating organizations operate diverse infrastructure—some may rely on Docker-hosted environments, others on Kubernetes clusters, while high-performance computing (HPC) centers frequently use Slurm for GPU scheduling. Previously, such differences often necessitated infrastructure standardization, adding friction to collaboration. NVIDIA FLARE circumvents this hurdle by adopting a two-layer architecture: persistent federation services for coordination and dynamic job execution tailored to each participant’s native environment.

With FLARE 2.9, federations can now seamlessly integrate Slurm-managed clusters, a critical addition for research institutions and enterprises leveraging HPC systems. Slurm joins Docker and Kubernetes as supported runtimes, unlocking flexibility while preserving local control over compute policies, datasets, and security settings.

Binance

How FLARE’s Architecture Works

FLARE separates long-running federation services from job execution. Persistent parent processes maintain the federation while job workers are launched dynamically to execute specific tasks. For example, when a data scientist submits an FL job, the system allocates required resources—GPUs, CPUs, and memory—based on study-specific configurations. Each site’s launchers translate these requirements into native execution units: Docker containers, Kubernetes pods, or Slurm batch jobs.

This design ensures resources are only consumed when jobs are actively running, minimizing overhead. It also allows organizations to maintain existing operational policies, such as node allocations in Kubernetes or quality-of-service rules in Slurm.

Market and Application Impact

The timing of FLARE 2.9’s release coincides with rapid growth in the federated learning market. According to Grand View Research, the sector is projected to reach $297.5 million by 2030, up from an estimated $167.1 million in 2026, representing a compound annual growth rate (CAGR) of 14.4%. This growth is driven by increasing demand for privacy-preserving AI solutions in sectors like healthcare, where organizations must comply with stringent data-sharing regulations while leveraging AI to analyze sensitive datasets.

By addressing infrastructure heterogeneity, FLARE makes federated learning more accessible for organizations that previously struggled with deployment complexity. Use cases include hospitals collaborating on medical image analyses, banks pooling fraud-detection models, or industrial players optimizing IoT systems. Each participant maintains control over its local resources while contributing to a shared model.

What’s Next?

With its expanded runtime support, NVIDIA FLARE is well-positioned to capitalize on federated learning’s growth trajectory. Organizations interested in deploying FLARE can explore its documentation or GitHub repository for implementation guidance. As the federated learning market evolves, interoperability and operational flexibility will remain critical, and NVIDIA’s approach could become a blueprint for scaling decentralized AI at production levels.

Image source: Shutterstock




Source link

Blockonomics

Be the first to comment

Leave a Reply

Your email address will not be published.


*