AI工具Score B (45)

Falcon-Emirati: When an LLM Learns the Dialect, the Culture, and the Nuance

20 小时前1 viewsSource: HuggingFace Blog

Falcon-Emirati: When an LLM Learns the Dialect, the Culture, and the Nuance

Team Article
Published October 6, 2026

image


Arabic is really a family of languages living under one name. Modern Standard Arabic is what you read in the news or a textbook, but it's rarely how people actually talk to each other. In the UAE, day-to-day conversation, humor, negotiation, and storytelling happen in Emirati Arabic, a Gulf dialect with its own vocabulary, its own rhythm, and a culture wrapped tightly around it. Emirati poetry, especially nabati poetry, along with proverbs and short anecdotes, carries meaning that doesn't survive a literal, word-for-word reading. A model that only knows MSA can translate every word of an Emirati sentence and still miss what it actually means.

That's the gap Falcon-Emirati-7B is built to close. It's a dialect-specialized model on top of Falcon-H1-Arabic, aimed at understanding and generating Emirati Arabic the way a native speaker would: the vocabulary, the tone, and the cultural context behind it.

Built on Falcon-H1-Arabic

We didn't start from scratch. Falcon-Emirati-7B is built on Falcon-H1-Arabic, our Arabic model family that already set new benchmarks for the language earlier this year. Falcon-H1-Arabic uses the Falcon-H1 hybrid architecture: State Space Models (Mamba) and Transformer attention running in parallel inside every block, with their outputs fused before each block's projection. That combination gives the linear-time efficiency of Mamba on long sequences while keeping the precision of attention for long-range dependencies, which matters for a morphologically rich language like Arabic. The family spans three scales (3B, 7B, and 34B parameters) with context windows up to 128K and 256K tokens, and it was already trained on a broad mix of MSA and dialectal Arabic (Gulf, Levantine, Egyptian, Maghrebi) alongside English and multilingual data.

That gave us a strong starting point: a model that already understood Arabic broadly, handled long context well, and had some dialectal exposure baked in. Falcon-Emirati-7B takes that foundation and pushes it specifically toward the Emirati dialect, the vocabulary, the grammar, and the cultural knowledge that a general Arabic model, however capable, doesn't pick up on its own.

We built Falcon-Emirati-7B on the 7B variant specifically. It's the sweet spot in the family: large enough to hold onto the nuance that dialect adaptation needs, but small enough that both training and inference stay practical. The 34B model would likely push quality a bit further, but at a training and serving cost that doesn't make sense for a dialect-specialized chat model, and the 3B model doesn't leave enough headroom for the depth of cultural and linguistic understanding we were after. 7B gave us the best balance of quality against training and inference cost.

Why Dialect Adaptation Is Hard

Turning a general Arabic model into an Emirati-dialect specialist sounds like a smaller job than building the base model in the first place. It isn't. A few things make it genuinely difficult:

  • Emirati is mostly a spoken dialect. It shows up far less in writing online than MSA, or even other Gulf and Levantine dialects, so there just isn't as much raw text to learn from.
  • Meaning is often non-literal. Idioms, proverbs, and poetic references lean on shared cultural context, not surface vocabulary.
  • There's no established playbook. There isn't a well-documented recipe for how much dialectal data is enough, how to mix it with MSA and general Arabic, or which training stage (continued pre-training, SFT, or preference optimization) matters most for picking up a dialect.

That last point shaped how we worked. A lot of building Falcon-Emirati-7B came down to trial and error: testing different data mixes, training stages, and supervision strategies, and using both human judgment and benchmark scores to figure out what actually moved the needle.

Our Approach to Data

We built a dedicated Emirati data pipeline on top of Falcon-H1-Arabic's pretraining, drawing on three complementary sources.

1. Authentic Emirati-Dialect Web Data

We crawled and curated content from Emirati websites and forums written natively in the dialect, not translated or transliterated from MSA. This is where we got our ground truth: how Emiratis actually write and speak online, the everyday phrasing, the colloquial expressions, and the natural back-and-forth between Emirati and MSA that shows up in real usage.

2. MSA Data About Emirati Culture and Identity

Alongside the dialectal text, we pulled in MSA-language material specifically about Emirati culture, heritage, and language: articles and references on local customs, values, history, and social norms, including how Emiratis are perceived and stereotyped. This doesn't teach the model to write in dialect, but it teaches the model what it's talking about when Emirati topics come up, things like heritage, etiquette, and the context a native speaker just knows.

3. Synthetic Data, Guided by Glossaries and Style Rules

Authentic dialectal text alone wasn't enough to cover the range of topics a chat model actually needs to handle day to day. So we generated a large amount of synthetic Emirati-dialect data to fill the gaps. We didn't just let a generator model improvise in "Gulf-ish" Arabic. We constrained it with strict rules and glossaries and dictionaries built specifically for Emirati vocabulary and grammar. Those guardrails made the difference between synthetic output that reads as authentically Emirati and output that's grammatically fine but sounds off to anyone who actually speaks the dialect.

Finding the Right Adaptation Recipe

Since there's no standard recipe for MSA-to-dialect adaptation, we treated the training strategy itself as something to figure out experimentally. We ran ablations on how much dialectal data to inject and at which stage of training, how to balance authentic crawled data against synthetic data without the model overfitting to synthetic patterns, and how much MSA cultural context was actually needed to keep it culturally grounded rather than just fluent on the surface. At each step we leaned on a mix of automatic scoring and native-speaker review, since automatic metrics alone don't capture naturalness, tone, or cultural fit well enough to trust on their own.

Evaluation Methodology

We tracked progress throughout training with two complementary approaches:

Manual Evaluation by Native Speakers

Emirati native speakers reviewed model outputs directly, judging not just whether an answer was correct but whether it sounded right: naturalness, tone, cultural appropriateness. These are the things a benchmark score won't tell you but a native ear catches immediately.

Automatic Evaluation on Alyah

For quantitative tracking, we used Alyah (الياه, "North Star"), a benchmark we and the community released specifically to evaluate Emirati-dialect capability in Arabic LLMs. Alyah is a fully native multiple-choice benchmark of 1,173 samples, collected manually from native Emirati speakers and spanning categories from everyday greetings and etiquette to figurative language, heritage knowledge, and Emirati poetry: the categories where dialect and culture matter most and where generic Arabic models tend to struggle. Full details on Alyah's construction and category breakdown are available in our benchmark blog post, and background on the base model family is available in the Falcon-H1-Arabic announcement.

Read the full original article:

HuggingFace Blog

#大模型#阿联酋