assemsabry commited on
Commit
39af9f3
·
verified ·
1 Parent(s): 3ed3889

Release Arabic-only Horus Taleeq 0.2 Base SLM under MIT

Browse files
LICENSE ADDED
@@ -0,0 +1,21 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ MIT License
2
+
3
+ Copyright (c) 2026 TokenAI and Assem Sabry
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
README.md ADDED
@@ -0,0 +1,145 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language:
3
+ - ar
4
+ tags:
5
+ - arabic
6
+ - arabic-language-model
7
+ - causal-lm
8
+ - text-generation
9
+ - egyptian-arabic
10
+ - arabic-dialects
11
+ - base-model
12
+ - slm
13
+ - horus
14
+ - tokenai
15
+ license: mit
16
+ pipeline_tag: text-generation
17
+ library_name: transformers
18
+ ---
19
+
20
+ # Horus Taleeq 0.2 Base
21
+
22
+ Horus Taleeq 0.2 Base is an **Arabic-only Small Language Model (SLM)** with a
23
+ decoder-only causal Transformer architecture, developed by TokenAI. It is a
24
+ ready-to-build-on base checkpoint, not a finished instruction-tuned or production
25
+ chat release.
26
+
27
+ ## Ownership and development
28
+
29
+ - Company: [TokenAI](https://tokenai.llc/)
30
+ - Owner and founder: **Assem Sabry**
31
+ - Lead developer: **Assem Sabry**
32
+ - Hugging Face organization: [tokenaii](https://huggingface.co/tokenaii)
33
+ - Model repository: [tokenaii/Horus-Taleeq-0.2-base](https://huggingface.co/tokenaii/Horus-Taleeq-0.2-base)
34
+
35
+ ## Model summary
36
+
37
+ | Item | Value |
38
+ | --- | --- |
39
+ | Model family | Horus Taleeq |
40
+ | Model type | Arabic-only Small Language Model (SLM), decoder-only causal Transformer |
41
+ | Approximate parameters | 0.2B (about 204.6M) |
42
+ | Primary language | Arabic |
43
+ | Intended text | Modern Standard Arabic and Arabic dialectal text |
44
+ | Egyptian Arabic | Targeted in the training program; not independently certified by this base checkpoint |
45
+ | Vocabulary | 128,000-token Arabic-first SentencePiece tokenizer |
46
+ | Hidden size | 640 |
47
+ | Transformer layers | 24 |
48
+ | Attention heads | 10 |
49
+ | Key/value heads | 2 (GQA) |
50
+ | MLP intermediate size | 2,048 |
51
+ | Activation | SiLU/SwiGLU-style Llama MLP |
52
+ | Normalization | RMSNorm, epsilon 1e-5 |
53
+ | Position encoding | RoPE, theta 1,000,000 |
54
+ | Embeddings | Tied input and output embeddings |
55
+ | Tensor type | bfloat16 |
56
+ | Context configuration | 4,096 tokens; initial pretraining sequences used 2,048 tokens |
57
+
58
+ The uploaded checkpoint is the completed `clean-repair-v1` base checkpoint. The
59
+ later canonical-clean continuation is still a separate development run and is
60
+ not included in this upload.
61
+
62
+ ## Training status
63
+
64
+ This repository is intended for the verified base checkpoint only. The original
65
+ training ledger records approximately 6.0B tokenizer IDs across earlier training
66
+ streams, but those historical streams are not yet a fully reproducible public
67
+ corpus. A separate canonical-clean continuation is being evaluated and must not
68
+ be represented as complete until its checkpoint and data manifest are published.
69
+
70
+ The current release therefore makes no claim of finished chat quality, factual
71
+ benchmark leadership, or production readiness. Earlier conversation adapters are
72
+ not part of this base repository.
73
+
74
+ The reproducibility baseline uses fused AdamW with betas `(0.9, 0.95)`, epsilon
75
+ `1e-8`, weight decay `0.1`, gradient clipping `1.0`, CUDA bfloat16 autocast, and
76
+ activation checkpointing. STAM is not the optimizer used for this verified base
77
+ checkpoint.
78
+
79
+ ## Intended use
80
+
81
+ - Arabic language-model research and evaluation
82
+ - Continued pretraining and supervised fine-tuning experiments
83
+ - Arabic and dialectal text generation research
84
+ - Building a separate instruction/chat checkpoint after evaluation
85
+
86
+ This base model is ready for continued pretraining, instruction tuning, identity
87
+ tuning, and downstream Arabic applications. It is not a safety-tuned assistant
88
+ and should not be used as a production chatbot without additional alignment,
89
+ safety, and quality evaluation.
90
+
91
+ ## Limitations
92
+
93
+ - Base-model completion quality may be inconsistent and may repeat text.
94
+ - It does not guarantee factual accuracy or reliable arithmetic.
95
+ - Dialect fluency and dialect identification require held-out evaluation.
96
+ - The model does not expose hidden chain-of-thought and should not be prompted
97
+ to reveal one.
98
+ - The historical training ledger and the later cleaned Parquet export are not
99
+ identical; users must not claim full corpus reproducibility from this release.
100
+
101
+ ## Files in this repository
102
+
103
+ The planned base release contains the model configuration, bfloat16 weights,
104
+ generation defaults, the matching 128K base tokenizer, the MIT license, and this
105
+ model card. The role-special tokenizer created for future chat SFT is deliberately
106
+ not presented as the tokenizer used by this base pretraining release. No SFT
107
+ adapter, teacher outputs, private credentials, or raw training corpus is included.
108
+
109
+ ## Loading with Transformers
110
+
111
+ ```python
112
+ from transformers import AutoModelForCausalLM, AutoTokenizer
113
+
114
+ repo = "tokenaii/Horus-Taleeq-0.2-base"
115
+ tokenizer = AutoTokenizer.from_pretrained(repo)
116
+ model = AutoModelForCausalLM.from_pretrained(
117
+ repo,
118
+ torch_dtype="auto",
119
+ device_map="auto",
120
+ )
121
+
122
+ prompt = "اكتب فقرة قصيرة عن أهمية القراءة."
123
+ inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
124
+ outputs = model.generate(**inputs, max_new_tokens=120, do_sample=False)
125
+ print(tokenizer.decode(outputs[0], skip_special_tokens=True))
126
+ ```
127
+
128
+ ## Citation
129
+
130
+ ```bibtex
131
+ @misc{tokenai_horus_taleeq_02_base_2026,
132
+ title = {Horus Taleeq 0.2 Base},
133
+ author = {Assem Sabry and TokenAI},
134
+ year = {2026},
135
+ publisher = {Hugging Face},
136
+ howpublished = {\\url{https://huggingface.co/tokenaii/Horus-Taleeq-0.2-base}},
137
+ note = {Arabic-first decoder-only base language model}
138
+ }
139
+ ```
140
+
141
+ ## License
142
+
143
+ This model is released under the [MIT License](LICENSE). Users remain responsible
144
+ for complying with applicable laws, third-party data rights, and the terms of any
145
+ data sources used in downstream training or evaluation.
config.json ADDED
@@ -0,0 +1,30 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "_name_or_path": "/mnt/opet-data/slm-nour-flash/checkpoints/horus-taleeq-ctx4096/final",
3
+ "architectures": [
4
+ "LlamaForCausalLM"
5
+ ],
6
+ "attention_bias": false,
7
+ "attention_dropout": 0.0,
8
+ "bos_token_id": 1,
9
+ "eos_token_id": 2,
10
+ "head_dim": 64,
11
+ "hidden_act": "silu",
12
+ "hidden_size": 640,
13
+ "initializer_range": 0.02,
14
+ "intermediate_size": 2048,
15
+ "max_position_embeddings": 4096,
16
+ "mlp_bias": false,
17
+ "model_type": "llama",
18
+ "num_attention_heads": 10,
19
+ "num_hidden_layers": 24,
20
+ "num_key_value_heads": 2,
21
+ "pretraining_tp": 1,
22
+ "rms_norm_eps": 1e-05,
23
+ "rope_scaling": null,
24
+ "rope_theta": 1000000.0,
25
+ "tie_word_embeddings": true,
26
+ "torch_dtype": "bfloat16",
27
+ "transformers_version": "4.46.3",
28
+ "use_cache": false,
29
+ "vocab_size": 128000
30
+ }
generation_config.json ADDED
@@ -0,0 +1,6 @@
 
 
 
 
 
 
 
1
+ {
2
+ "_from_model_config": true,
3
+ "bos_token_id": 1,
4
+ "eos_token_id": 2,
5
+ "transformers_version": "4.46.3"
6
+ }
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:5f20632c928185b007b5c86656aada7a588ebf6e7fecb09c9f47b7d6cd7e4621
3
+ size 399856864
special_tokens_map.json ADDED
@@ -0,0 +1,23 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "bos_token": {
3
+ "content": "<s>",
4
+ "lstrip": false,
5
+ "normalized": false,
6
+ "rstrip": false,
7
+ "single_word": false
8
+ },
9
+ "eos_token": {
10
+ "content": "</s>",
11
+ "lstrip": false,
12
+ "normalized": false,
13
+ "rstrip": false,
14
+ "single_word": false
15
+ },
16
+ "unk_token": {
17
+ "content": "<unk>",
18
+ "lstrip": false,
19
+ "normalized": false,
20
+ "rstrip": false,
21
+ "single_word": false
22
+ }
23
+ }
tokenizer.model ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:116147d034e7935a54e6a438fa078b390b0726eab14eb40829e3270c5b6f6fd4
3
+ size 2953471
tokenizer_config.json ADDED
@@ -0,0 +1,42 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_bos_token": false,
3
+ "add_eos_token": false,
4
+ "add_prefix_space": true,
5
+ "added_tokens_decoder": {
6
+ "0": {
7
+ "content": "<unk>",
8
+ "lstrip": false,
9
+ "normalized": false,
10
+ "rstrip": false,
11
+ "single_word": false,
12
+ "special": true
13
+ },
14
+ "1": {
15
+ "content": "<s>",
16
+ "lstrip": false,
17
+ "normalized": false,
18
+ "rstrip": false,
19
+ "single_word": false,
20
+ "special": true
21
+ },
22
+ "2": {
23
+ "content": "</s>",
24
+ "lstrip": false,
25
+ "normalized": false,
26
+ "rstrip": false,
27
+ "single_word": false,
28
+ "special": true
29
+ }
30
+ },
31
+ "bos_token": "<s>",
32
+ "clean_up_tokenization_spaces": false,
33
+ "eos_token": "</s>",
34
+ "legacy": false,
35
+ "model_max_length": 1000000000000000019884624838656,
36
+ "pad_token": null,
37
+ "sp_model_kwargs": {},
38
+ "spaces_between_special_tokens": false,
39
+ "tokenizer_class": "LlamaTokenizer",
40
+ "unk_token": "<unk>",
41
+ "use_default_system_prompt": false
42
+ }