IFM has released K2 Horizon, the largest and most transparent open-source model fleet in AI history, spanning from 0.9B to 375B parameters.

On September 3, 2026, the landscape of artificial intelligence shifted fundamentally. IFM has officially released K2 Horizon, a connected fleet of six foundation models that represents the largest and most significant open-source launch in the history of the field. For developers and AI engineers, this isn't just another model release; it is a declaration of transparency and a massive leap forward for the decentralized AI ecosystem.
The K2 Horizon project is designed to bridge the gap between massive, proprietary closed-source models and the highly efficient, specialized models used in edge computing. By providing a spectrum of intelligence, IFM allows engineers to deploy the exact level of reasoning required for their specific hardware constraints, without sacrificing the state-of-the-art performance that was previously locked behind expensive API walls.
Unlike monolithic releases, K2 Horizon is a 'connected fleet' of six distinct models. This architecture allows for seamless scaling across different computational environments. The fleet ranges from a lightweight 0.9B parameter model designed for ultra-low latency edge devices to a massive 375B parameter powerhouse capable of complex multi-step reasoning and high-level cognitive tasks.
What sets K2 Horizon apart is the radical transparency of its development. IFM has released not just the weights, but the full open code, the complete training data sets, and the precise training recipes used to achieve these results. This level of disclosure is unprecedented and provides a gold standard for reproducible AI research.
The performance metrics for K2 Horizon are nothing short of staggering. In the small-scale category, the 0.9B, 3.7B, and 7B models have officially set new State-of-the-Art (SOTA) benchmarks for their respective parameter counts. This means that a 7B K2 Horizon model can outperform many legacy 13B or even 30B models from previous generations in critical reasoning tasks.
When evaluating coding and agentic capabilities—the two most critical metrics for modern developers—K2 Horizon dominates. On HumanEval, the 7B model shows a significant delta over its competitors, while the 375B flagship model rivals the top-tier proprietary models in SWE-bench, demonstrating an uncanny ability to navigate complex software repositories and execute autonomous agentic tasks.
Because K2 Horizon is a fully open-source release, the primary mode of consumption for many will be local deployment via frameworks like vLLM or Ollama. However, for those who prefer managed infrastructure, IFM provides high-performance API endpoints designed for low-latency production environments.
While the models are free to download and host yourself, the managed API pricing is structured to be highly competitive with existing industry leaders, ensuring that developers can scale from prototyping to production without a massive spike in overhead.
The versatility of the K2 Horizon fleet makes it applicable to almost every vertical in the AI industry. The smaller models (0.9B to 7B) are perfect for on-device AI, such as real-time coding assistants in IDEs, privacy-focused local chat agents, and embedded systems that require immediate response times without cloud connectivity.
For enterprise-grade applications, the 120B and 375B models are the clear choice. These models excel in complex RAG (Retrieval-Augmented Generation) pipelines, multi-agent orchestration, and deep reasoning tasks where accuracy is non-negotiable. Whether you are building a sophisticated autonomous software engineer or a massive-scale knowledge management system, there is a K2 Horizon model tailored for the job.
Getting started with K2 Horizon is straightforward for anyone familiar with the modern AI stack. The model weights and training code are available via the official IFM GitHub repository and Hugging Face. For developers looking for a quick start, IFM provides a dedicated Python SDK that simplifies the process of loading the models and configuring inference parameters.
For production-grade deployments, we recommend using the K2-Containerized environment, which comes pre-configured with optimized kernels for both NVIDIA and specialized AI accelerators. This ensures you get the maximum throughput possible from the K2 architecture from day one.