The languages the web never wrote down.

Mouseion Labs licenses native-language source material for model training. We work with the institutions that hold a language’s written record — national libraries, university archives, newspaper houses — digitize it on site, and deliver the page images with a bibliographic record attached to every item. You run your own transcription.

We begin in Egypt, with the largest Arabic publishing record in existence.

1,351,921

Arabic-language items identified

Reported
79.7B

Estimated tokens after transcription

Estimate
4

Source institutions

Reported
1870

Continuous legal deposit since

Reported

Egypt corpus, revision 2026.09. Every figure is sourced on the provenance record.

Why the web is not enough

Model quality follows training data. For every language but English, the gap is a data gap before it is anything else.

English got there first

The breadth of English online is unmatched, and labs have gone further, digitizing English hardcopy to add what the web never carried.

Other languages run out of text

Labs crawl whatever native text exists online. For most languages that runs out quickly, and model quality follows it down.

Translation is a lossy substitute

When native text runs out, English is machine-translated into the target language. What comes back is English thinking in a different script; idiom, register and dialect are the first things lost.

The written record of most languages is on shelves, not servers. That is what we license.

What we do

Four steps, in order, for every collection we take on. Skip the manifest and you have images nobody can trace.

  1. 1

    Source

    We find the collections that were never online to be crawled — national legal deposit, university archives, century-long newspaper runs — and establish what is actually held, item by item, rather than relying on a published headline figure.

  2. 2

    License

    We negotiate directly with the institution that holds the material. The institution keeps ownership of its collection in every arrangement we make, and receives a share of what the corpus earns.

  3. 3

    Digitize

    Non-destructive imaging on site, so nothing in the collection is consumed to produce the corpus. Page images are captured at archival resolution, with a capture record for every page: item identifier, page order, capture date and resolution.

  4. 4

    Deliver

    Corpora ship as page images, with a bibliographic record for every item — title, author, imprint, date, shelf mark and the institution it came from. You can trace any page back to the shelf it came from.

What you receive

Every corpus ships in the same four parts, whatever the language or the collection it came from.

Page images

Archival masters in JPEG 2000 or TIFF, with access derivatives. Masters are kept permanently, so higher-resolution derivatives can be issued later.

Page sequence and capture record

Every image carries its item identifier, page order, capture date and resolution, so whatever you transcribe maps back to the page.

Bibliographic manifest

One record per item: title, author, imprint, place, date, shelf mark, holding institution and capture date. Delivered as JSON Lines or MARC.

License and chain of title

The agreement under which the material was digitized, naming the holding institution and the scope of use granted. Documented rather than asserted.

Where we work

Egypt is the first region, not the only one. We choose collections the same way each time: a language with far more speakers than written material available for training, and a record held somewhere that has never been digitized.

Egypt

Arabic

The largest continuous publishing record in the Arab world, held across four institutions and gathered under legal deposit since 1870. We are licensing the twentieth-century technical, scientific, legal and periodical record, along with the manuscript collections.

See the catalog

In development

Middle East

Arabic

Jordan and Saudi Arabia first, then the rest of the region's national and university libraries.

Scoped

East Africa

Swahili

More than a hundred million speakers, and a printed record out of all proportion to what exists online.

Scoped

Hold a collection?

If you are a library, a university system or a publisher with a backlist, we license rather than scrape. The institution keeps ownership of its material and shares in what the corpus earns.