The languages the web never wrote down.
Mouseion Labs licenses native-language source material for model training. We work with the institutions that hold a language’s written record — national libraries, university archives, newspaper houses — digitize it on site, and deliver the page images with a bibliographic record attached to every item. You run your own transcription.
We begin in Egypt, with the largest Arabic publishing record in existence.
Arabic-language items identified
Estimated tokens after transcription
Source institutions
Continuous legal deposit since
Egypt corpus, revision 2026.09. Every figure is sourced on the provenance record.
Why the web is not enough
Model quality follows training data. For every language but English, the gap is a data gap before it is anything else.
English got there first
The breadth of English online is unmatched, and labs have gone further, digitizing English hardcopy to add what the web never carried.
Other languages run out of text
Labs crawl whatever native text exists online. For most languages that runs out quickly, and model quality follows it down.
Translation is a lossy substitute
When native text runs out, English is machine-translated into the target language. What comes back is English thinking in a different script; idiom, register and dialect are the first things lost.
The written record of most languages is on shelves, not servers. That is what we license.
What we do
Four steps, in order, for every collection we take on. Skip the manifest and you have images nobody can trace.
- 1
Source
We find the collections that were never online to be crawled — national legal deposit, university archives, century-long newspaper runs — and establish what is actually held, item by item, rather than relying on a published headline figure.
- 2
License
We negotiate directly with the institution that holds the material. The institution keeps ownership of its collection in every arrangement we make, and receives a share of what the corpus earns.
- 3
Digitize
Non-destructive imaging on site, so nothing in the collection is consumed to produce the corpus. Page images are captured at archival resolution, with a capture record for every page: item identifier, page order, capture date and resolution.
- 4
Deliver
Corpora ship as page images, with a bibliographic record for every item — title, author, imprint, date, shelf mark and the institution it came from. You can trace any page back to the shelf it came from.
What you receive
Every corpus ships in the same four parts, whatever the language or the collection it came from.
Page images
Archival masters in JPEG 2000 or TIFF, with access derivatives. Masters are kept permanently, so higher-resolution derivatives can be issued later.
Page sequence and capture record
Every image carries its item identifier, page order, capture date and resolution, so whatever you transcribe maps back to the page.
Bibliographic manifest
One record per item: title, author, imprint, place, date, shelf mark, holding institution and capture date. Delivered as JSON Lines or MARC.
License and chain of title
The agreement under which the material was digitized, naming the holding institution and the scope of use granted. Documented rather than asserted.
Where we work
Egypt is the first region, not the only one. We choose collections the same way each time: a language with far more speakers than written material available for training, and a record held somewhere that has never been digitized.
Egypt
Arabic
The largest continuous publishing record in the Arab world, held across four institutions and gathered under legal deposit since 1870. We are licensing the twentieth-century technical, scientific, legal and periodical record, along with the manuscript collections.
Middle East
Arabic
Jordan and Saudi Arabia first, then the rest of the region's national and university libraries.
East Africa
Swahili
More than a hundred million speakers, and a printed record out of all proportion to what exists online.
Hold a collection?
If you are a library, a university system or a publisher with a backlist, we license rather than scrape. The institution keeps ownership of its material and shares in what the corpus earns.