Authors
Safouane Boudakkou, Abdelaaziz El Hibaoui
Published in
Data in brief. Volume 68. Pages 113140. Epub Aug 05, 2026.
Abstract
This article presents AryWiki-Instruct, a high-fidelity instruction tuning dataset for Moroccan Arabic (Darija), comprising 46,590 Question and Answer (QA) pairs. The dataset was derived from a snapshot of the Moroccan Arabic Wikipedia (arywiki) and generated using the Gemini-2.5-Flash model via a Context Aware batch processing architecture. The data creation process involved parsing raw XML Wikipedia dumps, filtering for script consistency and token density, and applying a Context Injection generation strategy to prevent coreference ambiguity. To ensure high information density, the raw generated output (82,500 pairs) was subjected to a rigorous automated quality assurance pipeline. This pipeline utilized a hierarchical trigger confirmation algorithm to remove repetitive administrative census noise and employed composite embedding based clustering (multilingual-e5-large) for semantic deduplication. The final dataset is formatted as a JSONL file, providing paired instructions and responses alongside their taxonomic categories. This dataset provides a native first culturally grounded resource designed to facilitate the supervised fine tuning of Large Language Models (LLMs) in Maghrebi dialects, circumventing the syntactic limitations of translated English instruction sets.
PMID:
42633346
Bibliographic data and abstract were imported from PubMed on 23 Aug 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 1
- Comments 0