Introducing Falcon-H1-Arabic: Pushing the Boundaries of Arabic Language AI with Hybrid Architecture

Falcon-H1-Arabic is a new family of advanced Arabic language models featuring a hybrid Mamba-Transformer architecture and support for massive context windows up to 256K tokens.
The journey of building world-class Arabic language models has been one of continuous learning and iteration. Today, we're excited to announce Falcon-H1-Arabic, our most advanced Arabic language model family to date, representing a significant leap forward in both architecture and capabilities. This release embodies months of research, community feedback, and technical innovation, culminating in three powerful models that set new standards for Arabic natural language processing. Falcon-H1-Arabic is built on the Falcon-H1 hybrid architecture, which integrates State Space Models (Mamba) and Transformer attention within every block. This design provides the linear-time scalability of Mamba for extremely long sequences while preserving the precise long-range modeling capabilities of attention. We've dramatically increased context capabilities from 32K to 128K tokens for the 3B model and 256K tokens for both the 7B and 34B models. We rebuilt our pre-training data pipeline from the ground up to better reflect the complexity of Arabic, including dialect coverage for Egyptian, Levantine, Gulf, and Maghrebi. After pre-training, the models undergo a focused post-training pipeline consisting of supervised fine-tuning (SFT) followed by direct preference optimization (DPO). On the Open Arabic LLM Leaderboard (OALL), Falcon-H1-Arabic achieves state-of-the-art results at every scale tested, with the 3B model outperforming similar-sized competitors by significant margins.
Source: Hugging Face Blog
















