Markdown that still has
the document in it.

Kerf reads PowerPoint decks and live web pages. Tables come back as tables. Code comes back as code. Headings arrive in the order the page put them, and the cookie banner stays where you found it.

Get started for free

Get 250 credits free when you link a card

A pricing page in, Markdown out.

captured 2026-09-12

Markdown comes back here.

200 saved14,696 chars1 table55 blocks of furniture left out0/153

The address above is a saved answer from a real call, replayed a paragraph at a time: the links, the image lines, the text the page keeps folded away behind a click, and the whole eighty-row comparison table. Type any other address and Kerf reads it live. The first five pages are free, and the demo follows each site's robots.txt.

76/76

pages came back

28/29

tables stayed tables

90%

of headings kept

587ms

for a typical page

/ What you get

Extraction usually costs you the parts of a document that carried the meaning.

A table becomes a paragraph of numbers. A heading disappears. A slide arrives shuffled. Kerf was built to stop that, and here is what that is worth on a fixed set of real pages.

28/29

Pages whose source holds a table

A table you can still ask questions about

Pricing tiers, rate limits, a comparison grid. Flatten one into a paragraph and every number in it loses the row and column that gave it meaning. Kerf brought back 28 of the 29 tables in our test set. The provider we ran beside it brought back 8.

90%

Share of the source page's headings that arrived

Headings your index can actually use

A dropped heading is a section your retrieval cannot cite and your reader cannot find. 90% of the page's own headings came through, against 76% beside us.

587ms

Median wall clock, as the caller sees it

Fast enough to sit inside a request

Most pages never need a browser, so Kerf reads them straight from the markup and only spins one up for the pages that genuinely need it. That is about 2.9 times quicker than the alternative managed.

/ Web pages

Send a URL. Get the page back, without the furniture around it.

Kerf Online renders the page the way a browser does, then returns the document. Navigation, cookie banners, newsletter boxes and footers are left out. What survives is what somebody wrote.

A table is a grid before it is text, and the cut recognises the grid before it reads a word of it. Flattened into prose, a pricing table stops being answerable.

Kerf
28/29
TinyFish
8/29

Pages whose source holds a table · 76 fixed addresses · our own run

Most pages never need a browser. Kerf reads them straight from the markup and only spins one up for the pages that genuinely need it, which is roughly a third of the time the alternative took.

Kerf
587ms
TinyFish
1,724ms

Median wall clock as the caller sees it · the bar is the time taken · our own run

/ Decks

A slide deck is the worst-behaved document your company owns.

PowerPoint stores shapes in the order somebody drew them, which is almost never the order anybody reads them. Ask a model about the deck below and shape order puts it on the middle card.

E_grouped_cards.pptx · 8 cuts
A slide headed Grouped cards, with three cards side by side: Swim, Wear and Pair, each with a subtitle and a sentence beneath it.

#### Grouped cards

Swim

Built for water

Airless link keeps audio flowing where Bluetooth stops.

Wear

Swim-first headset

A low-profile headset that sits under cap and goggles.

Pair

Just your phone

Pairs with your phone like any normal earphones.

Kerf reads down each column, because the blank gutters between the cards say there are three of them.

/ The numbers

We called Kerf and TinyFish on the same 76 addresses and saved every answer before scoring any of it.

The addresses were fixed before either was called, and picked to still mean the same thing next year: versioned docs, published standards, licences, essays nobody edits. Each answer is scored against the page's own HTML. No model marks this homework, so anyone can run it again and get the same result.

Pages returned

Kerf
76/76
TinyFish
75/76

At least 1,000 characters came back, which is the floor TinyFish set in their own evaluation.

Headings kept

Kerf
90%
TinyFish
76%

Share of the page's own headings that arrived at all. A dropped heading is a section an agent cannot cite.

Headings in order

Level
Kerf
100%
TinyFish
100%

Every pair of headings the page put in an order, still in that order. Neither of us gets this wrong on the web.

Tables kept as tables

Kerf
28/29
TinyFish
8/29

Of the pages whose source holds a table.

Code kept as code

Kerf
38/38
TinyFish
36/38

Of the pages whose source holds a code block. Indentation and line breaks are the meaning in a code sample.

Median characters

Kerf
19,654
TinyFish
18,008

More of the document came back, at the same measured level of boilerplate.

Median time a page

Less is better
Kerf
587ms
TinyFish
1,724ms

Wall clock, as the caller sees it. The bar is the time taken, so the short one is the quick one.

Pages returned, by kind of page

Documentation

10/1010/10

Reference

8/88/8

Standards

8/88/8

Pricing

8/88/8

Repositories

8/88/8

Longform

12/1212/12

Spa

8/88/8

Tables

6/65/6

Legal

6/66/6

Simple

2/22/2

Kerf first, TinyFish second. Another 5 addresses sit behind a bot challenge and are counted on their own. A site that says no has said no, and we answer that rather than working around it.

Read the whole run

Every address, every result, what was dropped and why, and the commands to do it again.

/ What it reads

One reader, two formats today and a third on the way.

Decks

.pptx

Three cards across a slide come back as three cards, in the order you would read them. Speaker notes and alt text come with them. Feed a quarterly review to a model and it answers about the right quarter.

Running inside Quorum today, behind document storage and Attest.

Web pages

any URL

Send a URL, get the page. Pricing grids stay grids, API docs keep their code samples, and the nav, the footer and the newsletter box do not come along.

On the Enterprise API, and benchmarked against the field.

Scans and images

soon

photographed and scanned pages

A scanned contract or a photographed page is the hardest version of the same problem, and the one Kerf was built for first. Columns have to be found before a word can be read.

In progress.

/ Pricing

A tenth of a cent a page, and nothing for a page that does not come back.

For developers

$0.001

a web page

A tenth of a credit for every page Kerf reads, whether it takes the fast path or needs a browser. A credit is one US cent.

Free

when a page fails

A site that refuses and a page that will not load cost nothing. A page served again from cache is free too.

2,500 pages

to start

250 credits, granted once to a developer workspace that links a card, with nothing charged to it.

For teams

If Kerf belongs inside a pipeline your organisation runs rather than a script you call, we will go through the deployment with you: the volumes, the sites you read, and how the pages reach your systems.

/ Start

One key, and Kerf is on it.

Kerf sits on the Quorum Enterprise API next to Attest, metered in credits like everything else there. One account, one invoice, and the key you already hold reaches both.

curl https://api.quorumtech.ch/kerf/v1/render \
  -H "Authorization: Bearer $QUORUM_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com/report"}'
POST /kerf/v1/render
kerf 0.4.0

The capture in the hero was taken on 2026-09-12 and every box, cut and line of Markdown on this page comes from it. The cut is a recursive XY-cut, the method document layout analysis has used to find columns since the 1990s; Kerf's name is the slot a saw blade leaves.

The head-to-head was run on 2026-09-12 over 76 addresses fixed before either provider was called, chosen to still resolve to the same document next year. Both providers were called by us through their own APIs, one after another, with no tuning and no retries, and every response was saved before anything was scored. Scoring compares each response against the page's own HTML, so reading order and structure are measured rather than judged.

A further 5 addresses sit behind a bot challenge and are reported separately, never as an extraction failure. Kerf waits out a wall that lifts by itself and names one that needs a person. It does not solve challenges, and a site that says no is answered as a site that said no.

The benchmark run asked Kerf to skip its robots.txt check so that extraction was measured against extraction rather than against another provider's compliance. Honouring robots.txt is the default everywhere else.

Kerf is part of the Quorum Enterprise API at developer.quorumtech.ch.