The platform

From raw text to findings you can defend.

Point Dovre at your corpus and it returns named themes with exact counts, how each one moves over time, and the messages behind every finding — so a result can be checked rather than taken on trust. Every document is read; nothing is sampled.

  • 10M+Rows per corpus
  • 100%Of documents read, never sampled
  • 9Languages clustered today
  • 7Source formats in
Inputs

Seven ways in.

A CSV export, a PDF, a live database table — each streamed in, so corpus size is a storage question rather than a memory ceiling.

CSV

Map a text column and an optional timestamp column. Streamed in batches, never loaded whole.

JSON

Records or line-delimited JSON, with the text field selected at ingest time.

Plain text

One document per file or per line, whichever matches how your corpus is stored.

PDF

Text extracted per page or per paragraph, so a long report becomes many documents.

Word (.docx)

Same paragraph-level extraction as PDF, for anything that started life in Word.

Postgres

Point at a table, pick the text column, and pull directly — no export step.

SQLite

The same SQL path for local or embedded databases.

Output

A topic is more than a label.

Every run returns the evidence behind each cluster, so a topic can be checked rather than taken on trust. This is what the dashboard shows for one — illustrative data, drawn in the same shapes your corpus would fill.

Topics & keywords

Every cluster named, with the c-TF-IDF keywords that define it ranked by importance.

Hierarchy

A dendrogram showing how topics merge, so you can read the corpus at any level of detail.

Volume over time

A trend line per topic — see what is growing, what is fading, and when it turned.

Representative docs

The documents closest to each cluster centre, so a topic is never just a label.

LLM insights

Optional per-topic naming and a written insight, generated from the cluster's own documents.

Classified coverage

The share of your corpus a run actually assigned, surfaced up front rather than buried.

Languages

One inbox, five languages, one topic.

Southeast Asian text does not arrive sorted by language, so the model does not ask it to be. Documents are clustered on meaning, which is language-independent — a complaint in Bahasa and the same complaint in Filipino land in the same topic, with no translation step. Keyword extraction is the part that is still catching up.

  • English
  • Bahasa Indonesia
  • Bahasa Melayu
  • Filipino

Clustering and topic naming work in these too — keyword extraction is in progress:

  • Tiếng Việt
  • ไทย
  • 中文
  • ខ្មែរ
  • မြန်မာ
Built to last

Built for real corpora, running today.

Built for 10M+ rows

Corpus size is bounded by disk, not memory, so ten-million-row runs are routine rather than the edge of what fits.

Streaming ingestion

Sources stream in as batches, so a run starts working before the whole corpus has landed — and is never held in memory at once.

Runs you can watch

Every run is observable end to end, with live progress from the first document to the finished topics — so a long job is accountable, not a black box.

Running today, not a roadmap

Real corpora go in and named findings come out on infrastructure that is already deployed — a product you can run this week.

Run it on your own corpus.

You bring the text and the question. Everything on this page is machinery we operate on your behalf — there is nothing here for you to install.