Life sciences · Preprint
arXiv · September 3, 2026
Early or partial results. Treat as a signal, not a conclusion.
This preprint describes curation and release of two new Armenian-language datasets (ArmWeb and ArmSTEM) and continued pretraining of Gemma-4-E4B to produce arm-gemma-e4b, the first open Armenian LLM with reproducible recipe. Ablation experiments show that news-only pretraining improves fluency at the cost of knowledge, and that verified STEM translation data recovers the knowledge loss. The work has not been peer reviewed.
Resource development and model training study with ablation experiments. Armenian language resources and models; dataset validation included human evaluation.. Intervention: Continued pretraining of Gemma-4-E4B on ArmWeb and ArmSTEM datasets.. Compared with: Existing open Armenian models and unadapted Gemma-4-E4B base model..
ArmWeb corpus comprises 4.37M validated Armenian news documents. ArmSTEM contains 373K parallel English-Armenian math and science problems with step-by-step solutions. arm-gemma-e4b outperforms every existing open Armenian model and its unadapted base.
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
A resource development and model release study with ablation experiments on a low-resource language; demonstrates feasibility and comparative improvement but lacks external validation, peer review, and clinical or deployed endpoints.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Pretraining data for Armenian, a morphologically rich and low-resource language, is scarce, and no open Armenian LLM has been released with the data and recipe needed to reproduce it. To address this gap, we curate and release two datasets. ArmWeb is an extensively validated corpus of 4.37M Armenian news documents. ArmSTEM is a parallel English-Armenian collection of 373K math and science problems with step-by-step solutions, translated into Armenian and verified through both answer-preserving LLM judgment and human evaluation. Continued pretraining of Gemma-4-E4B on these datasets yields arm-gemma-e4b, which outperforms every existing open Armenian model as well as its unadapted base, and is the first open Armenian LLM with complete training data and recipe. Our ablations show that news-only continued pretraining improves fluency while eroding knowledge, a pattern we also observe in existing Armenian models, and that a small share of verified translated STEM data reverses the loss. We further find that the largest public Armenian corpora overlap web-derived evaluation panels heavily, including a train/test self-overlap inside FineWeb-2. We openly release all data, models, and code.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.