For CX, research and analyst teams

Findings you can defend, from more text than anyone could read.

Dovre reads every document in your corpus — no sampling — and returns named themes with exact counts, how each moves over time, and the source text behind every one. So a finding can be checked, not taken on trust.

Why not just paste it into ChatGPT?

A summary is an opinion. A Dovre run is a count you can check.

General assistants are good at reading a handful of documents closely. They are not built to tell you how many of your 200,000 tickets are about one thing, or whether it is growing.

  • Where your text goes

    Pasting into ChatGPT or Claude

    The whole corpus goes to a third-party API to get an answer at all.

    A Dovre run

    Your corpus never reaches a third-party model API. It is processed on infrastructure we run, and on the default profile the modelling touches no network at all. Topic naming is the one step that calls an LLM, it is optional, and it sees a topic's keywords and at most six short excerpts — never the corpus.

  • How much text

    Pasting into ChatGPT or Claude

    A few hundred documents, once you have trimmed them to fit. Past that you are deciding what to leave out before you have read any of it.

    A Dovre run

    200,000 documents in a single run on the default settings, streamed in from Parquet. The ceiling is a config value you own, not a property of the format.

  • Coverage

    Pasting into ChatGPT or Claude

    You get an impression of whatever the model read. Nothing tells you what it skimmed, dropped, or never reached.

    A Dovre run

    Every document in the run is embedded and assigned. The share that actually landed in a topic is reported as a number, up front, rather than left for you to assume.

  • Counts you can use

    Pasting into ChatGPT or Claude

    “Several users mention billing.” Several out of how many? Enough to staff for, or four people having a bad week?

    A Dovre run

    Every topic carries an exact document count and its share of the corpus, so a theme can be ranked, tracked and argued about with numbers.

  • Change over time

    Pasting into ChatGPT or Claude

    You would paste each month in separately and hope the summaries are comparable enough to line up.

    A Dovre run

    Each topic gets a volume trend built from your own timestamps, bucketed by day, week, month or quarter depending on the span.

  • Same answer twice

    Pasting into ChatGPT or Claude

    Ask again next month and you get a fresh interpretation. Nothing tells you whether the theme moved or the wording did.

    A Dovre run

    On the default profile the same corpus and config produce the same topics — the hashing embedder is deterministic and the reducer and clusterer are seeded. That is what makes month-over-month tracking mean anything.

  • Evidence

    Pasting into ChatGPT or Claude

    You cannot click a sentence in a summary to see which tickets produced it. You take the summary on trust.

    A Dovre run

    Each topic keeps the documents closest to its centre, so every theme opens onto the real text that produced it.

Bring text in from anywhere.

Ingest streams in batches, so a source is never loaded whole — millions of rows in without filling memory.

CSV & text

Bulk-ingest millions of rows. Map your text and date columns; we stream it in.

PDF & Word

Drop in documents — we extract text per page or paragraph before modelling.

Databases

Point Dovre at a Postgres or SQLite table and pull a text column directly.

Plug-in connectors

New sources are drop-in modules — add or remove them without touching the core.

FAQ

Questions we get asked.

Is this BERTopic with a UI on top?

No. The pipeline is our own: seven stages — ingest, preprocess, embed, reduce, cluster, represent, name — each resolved from a component registry by config. You can swap the embedder, reducer or clusterer independently, which is the part an off-the-shelf pipeline does not give you.

Why not just paste it into ChatGPT or Claude?

For twenty documents, do. The difference shows up at volume: a chat window gives you an impression of what it read, with no count behind it and no way to tell what it skipped. Dovre embeds and assigns every document in the run, so each topic carries an exact document count, a share of the corpus, a volume trend and the documents nearest its centre. It is the difference between a summary and a census.

How large a corpus can it handle?

Ingest streams to Parquet in batches, so pulling in millions of rows never loads the source whole. A modelling run then works on up to 200,000 documents on the default settings — that is the `max_docs_in_memory` cap, and raising it trades RAM for corpus size. We would rather publish the real ceiling than a rounder one; it is already far past what fits in a chat context.

Will I get the same topics if I run it again?

Yes, on the default profile. The hashing embedder is deterministic by construction, and the reducer and clusterer are seeded, so the same corpus and the same config produce the same topics. That matters if you intend to track a theme from one month to the next rather than re-interpreting a fresh answer each time.

Do I need a GPU?

No — and you do not need to think about one. We run the infrastructure: the embedding service, the CUDA hosts behind it, the worker fleet and the storage. What that buys you concretely is the Quality profile — BGE-M3 embeddings, UMAP and HDBSCAN — at throughput, without anyone on your side provisioning a GPU box, sizing a batch or keeping a model server alive between runs. (For the curious: the Fast profile genuinely needs nothing but NumPy. That is an engineering fact about the engine, not a hardware bill we are passing to you.)

So what am I actually operating?

Nothing. You point Dovre at a source, confirm which column is the text and which is the timestamp, and read the result. There is no cluster to run, no model to download, no pipeline YAML to own, and no MLOps hire between you and a topic map. The seven-stage engine is swappable by config because that is how we tune your runs — not a configuration burden we are handing over.

Why is this possible now and not five years ago?

Two things changed. Multilingual embedding models got good enough that documents about the same thing land together whatever language they are written in — no translation step, no per-language pipeline, which is what made a mixed-language Southeast Asian inbox tractable at all. And every team now has an LLM that will happily read their text and give them an impression of it, which has made the absence of a count much more obvious than it used to be. Topic modelling is not new. Topic modelling that survives contact with five languages and 200,000 documents is.

What does it cost?

We scope it per engagement rather than publishing a price list, because the two things that drive cost — how much text you have and where the run has to happen — vary more between customers than anything else. A single-tenant deployment inside your own network is a different proposition from a corpus we can run on our infrastructure. Tell us the corpus size and the constraint and we will come back with a number.

Bring the text and the question.

Bring the text and the question. The GPUs, the embedding service, the pipeline and the storage are ours to run — there is nothing for you to provision, and no one to hire before you get an answer.