Bayaga.
A Hebrew and Aramaic language model, built from scratch.
Our own tokenizer, our own model, our own corpus of the classical Jewish canon, with no dependency on any existing language model at any stage. It runs live. The deep work we take on for clients.
בְּרֵאשִׁית בָּרָא אֱלֹהִים
A base model, trained from scratch. It continues text. Instruction tuning is next.
Own tokenizer · own model · own corpus
A language model for the classical Jewish canon, and every layer is ours.
Bayaga reads and writes classical Hebrew and Aramaic: Tanakh, Mishnah, Gemara, Rashi, Tosafot, Midrash, Halakhah, responsa. It runs live. You give it the start of a passage and it continues it in the right register, in the language of the text itself. Our own tokenizer, our own model, our own corpus, with no dependency on any existing language model at any stage.
Continues Hebrew and Aramaic passages in register, across the entire canon.
Runs live in the browser, from a command line, and as a tool for other agents.
Reads Hebrew and Aramaic with one tokenizer, the two languages interleaved as they are in the Talmud.
No borrowed model. No borrowed tokenizer. No borrowed data pipeline.
No dependency on any existing language model at any stage, from the first token to the live server. Most teams fine tune someone else’s model and call it their own. We built the tokenizer, the model and the corpus, because for a language the big labs treat as an afterthought, borrowing is where quality dies.
A root aware tokenizer for Semitic text
Hebrew and Aramaic carry meaning in a consonantal root that survives across every inflection. Off the shelf tokenizers shatter that root into fragments. Ours is morphological and lossless: it keeps a single root recognisable across all its forms, handles Hebrew and Aramaic together, and reconstructs the original text exactly, character for character. It is the subject of a published technical report.
A transformer trained from the first token
A roughly 134 million parameter transformer, pretrained from scratch on next token prediction over the canon, built with modern primitives: rotary position embeddings, RMSNorm, SwiGLU, grouped query attention, tied embeddings. Trained into existence on our own corpus, not adapted from a base checkpoint.
A canon rebuilt into training data
Thousands of curated sources cleaned into hundreds of millions of tokens, genre balanced across Talmud, Halakhah, Tanakh, Mishnah, responsa and more, with train, validation and test splits kept strictly disjoint. The next generation rebuilds this into billions of tokens across streams that align every base text with its commentaries and its cross references.
It beats the big labs on the language it was built for.
The tokenizer results are published in a public technical report, measured head to head on the same text.
How much less Hebrew costs to tokenize than under the tokenizer behind ChatGPT, head to head on the same text.
Meaning bearing roots preserved as a single token, measured against the gold standard reference.
Tokenizer, model and corpus, all built from scratch. No existing language model anywhere in the lineage.
The classical canon rebuilt into an aligned corpus, every base text tied to its commentaries and cross references.
A model shaped like the way the texts actually argue.
The Talmud is not prose to be predicted, it is a debate: a question, a difficulty, a resolution, a claim weighed against a counter claim across centuries. Bayaga is the first step of a longer thesis that treats that structure as first class, rather than flattening it into plain text.
Morphology first
A tokenizer that keeps the root visible is the foundation everything else stands on. Live and published.
Dialectical structure
Representing a passage as its argument, question, difficulty, resolution, so the model reasons with the dialectic instead of only imitating its surface.
Epistemic status
Tracking whether a claim is settled, disputed or unresolved, so confidence is a property the model carries rather than a tone it fakes.
If we can do this for a two thousand year old canon, we can do it for your domain.
Building a competent model for a language or a domain the big labs ignore is the work we take on for clients. If your problem lives where the general models are weak, this is the team that builds the specialist.
A model trained from scratch
For a domain, a language or a corpus where a general model falls short and fine tuning one is not enough either. We take it from raw data to a running model.
A tokenizer for your language
When the standard tokenizers waste your text, garble your script or shatter your morphology, we design one that fits the language you actually work in.
Fine tuning and alignment
Supervised fine tuning, preference tuning and instruction following on top of a base model, ours or yours, with the evaluation to prove it worked.
Evaluation that means something
Harnesses that measure a model on its own held out data, with metrics that track real quality, so you know what you have before it ships.
Novel model conception
Original architecture for problems the standard transformer recipe does not answer, taken from a thesis to a trained model.
Productionised inference
From checkpoint to a served model with a clean API, streaming, and clients for the browser, the command line and other agents.
Straight answers.
Bayaga is a language model for classical Hebrew and Aramaic that La Boetie built entirely from scratch: our own morphological tokenizer, our own transformer, our own corpus of the classical Jewish canon, with no dependency on any existing language model at any stage. It runs live as a base model that continues text in register, and it shows the model building we do for clients.
Where the general models are weak, we build the specialist.
A model from scratch, a tokenizer for your language, fine tuning, alignment, evaluation. Tell us the problem the big labs ignore, and we will tell you what the specialist looks like.