1. True Open Science vs. "Open Weights"
In recent years, many major AI labs have labeled their models as "open source" while merely releasing compiled model weights without the underlying training data or recipe pipelines. K2 Horizon aims to make the training process substantially more reproducible by publishing the underlying artifacts alongside the models.
IFM presents K2 Horizon as an open-science effort designed so researchers can inspect how the models were built, reproduce experiments and build on the released artifacts.
Researchers can now trace the evolution of neural representations across intermediate training checkpoints, giving the global community unprecedented visibility into the inner mechanisms of frontier AI.
2. The 6-Model Lineup: Edge Wearables to Enterprise Clouds
K2 Horizon spans 6 distinct model sizes tailored for specific compute budgets:
| Model Tier | Target Environment | Active Parameters | Specialized Capabilities |
|---|---|---|---|
| K2 Horizon 0.9B | Smartwatches & AR Glasses | 0.9B | Top-tier math reasoning & tool calling for micro on-device constraints. |
| K2 Horizon 3.7B | On-Device & Fine-Tuning | 3.7B | Industry-leading reasoning benchmarks in the sub-4B category. |
| K2 Horizon 7B | Smartphones & Edge PCs | 7B | Full software engineering and coding synthesis directly on mobile silicon. |
| K2 Horizon 32B | Local Servers & Laptops | 32B Dense | High-density reasoning designed for high-end workstations and private cloud servers. |
| K2 Horizon 36B | Enterprise Compute Efficiency | 4B Active (MoVA) | Delivers massive-scale performance using only 4B active compute parameters. |
| K2 Horizon 375B | Enterprise Agent Swarms | 23B Active (MoVA) | Flagship agent orchestrator for complex multi-step reasoning and enterprise workloads. |
3. Key Architectural Innovations
1. Diffusion Distillation (3x Faster Generation)
Traditional Large Language Models generate text sequentially—one token at a time—which creates compute bottlenecks during long outputs. K2 Horizon replaces pure autoregressive decoding with Diffusion Distillation. By compressing diffusion generative blocks into distilled inference steps, K2 Horizon generates entire token clusters in parallel, with IFM reporting roughly 3x faster generation in its announced approach.
2. Mixture of Value Attention (MoVA)
Standard attention mechanisms allocate identical compute resources across all tokens in a prompt. K2 Horizon introduces MoVA (Mixture of Value Attention), which routes calculations selectively through expert value pathways. This allows the 36B model to execute with only 4B active parameters, slashing VRAM footprint and operational energy costs.
4. Open Availability & Ecosystem Support
The complete K2 Horizon family is released under the permissive Apache 2.0 license for broad commercial and academic use, subject to the license terms.
- Hugging Face Repository: Full FP16 weights, GGUF quantized models for Ollama / llama.cpp, and complete training datasets.
- Inference Ecosystem: First-day runtime support across vLLM, SGLang, Cerebras, Nebius, and GGUF runtimes.
5. Official IFM Announcement on X
IFM also announced the K2 Horizon release on X. The embedded post below links directly to the institute's announcement.
View the official K2 Horizon announcement from IFM on X