Kerf Online against Firecrawl, on 639 RAG questions
We read the same pages with both products, cut each result into chunks the way a retrieval pipeline does, asked the same questions of the chunks a retriever picked from each, and had the answers written and graded by models that never knew which extraction they were reading. Here is the whole run: why a retrieval test is a different measure, what we compared, how it was scored, every page, and what the numbers do not show.
Quorum Technologies · Firecrawl fetched 13 September 2026 · Kerf fetched 14 September 2026 · scored 14 September 2026
01
A retriever never sees the page. It sees a few chunks of it.
Our other study scores Markdown against the page's own HTML: headings kept, tables kept, reading order. This one asks the question that study cannot: once the Markdown is cut up and indexed, does the right answer come back?
Almost nobody hands a whole page to a model. A retrieval pipeline cuts the Markdown into chunks of a few hundred to a couple of thousand characters, stores them, and at question time hands the model the few that look most like the question. What the model can answer depends on whether the fact it needs landed in one chunk, with enough around it to say what it is.
That is a property of the extraction as much as of the chunker. A price that arrives on its own line, three headings away from the name of the product, is retrievable only by luck. A table row that arrives without its column names answers nothing. A page that came back as a shell, because its figures are written in by a script the extractor never ran, answers nothing at all, however tidy the Markdown looks.
Why not measure this with the structure scores we already publish? Because they measure the wrong end of the pipe for this question. Headings kept is a good proxy for a person reading a document and a poor one for a retriever, which does not read headings, and the same page can score well on structure and badly on answers. So this study puts a retriever between the extraction and the reader and counts answers.
02
What we compared
Two products that turn a URL into Markdown, each called through its public API from the United States, and Kerf a second time with the option it offers for exactly this use.
| Column | Called as | What it is |
|---|---|---|
| Kerf | mode auto · location us | Ours, as a customer calls it: the direct read where a page allows it, a browser where it does not, and the routing between the two left to Kerf. Links off, which is what a retrieval caller sends. |
| Kerf, optimiseForRAG | + optimiseForRAG true | The same fetches, with the option that shapes the Markdown for a chunker: fewer headings, table rows that carry their column names, a product's name kept beside its figures. The same words come back. |
| Firecrawl | v2 scrape · location US | Browser-backed scraping, Markdown format, at its own defaults with the location set to match ours. Described here as they describe themselves. |
We picked Firecrawl because it is the extractor most retrieval pipelines reach for first, and because it fetches from the United States by default, which is where prices and terms on most of these pages are decided. Kerf was asked to fetch from there too, so neither column is reading a different country's page.
03
The test set
110 pages of the kind people put into retrieval systems, and 639 questions written from them.
Pricing pages, product pages, API references, standards, encyclopaedia articles, repositories and essays: pages with facts in them that someone would ask a model about. Up to six questions were written for each page by a model that could see two extractions of it, in random order and without knowing which product made which, and was told to draw half its facts from each. Every question names a short gold answer.
| Kind | n | What it asks for |
|---|---|---|
| numeric | 181 | A figure: a price, a limit, a version, a count. |
| fact | 139 | A name, a date, a definition, a yes or no. |
| list | 101 | Several items, all of which have to arrive. |
| qualified | 218 | The right answer depends on a qualifier the page sets: which plan, which model, which case. An answer lifted from the wrong part of the page is wrong. |
Kinds of question, and how many of each were scored.
Were the questions written to favour either product? The writer could not tell the products apart, and was told to take half its facts from each extraction. The audit below checked every gold answer it could against the page.
04
How we score
The same chunker and the same retriever for every extraction, so any difference in what the answering model sees comes from the extraction and nothing else. Nobody in the chain knows which product they are reading.
| Step | What happens |
|---|---|
| Chunk | Each extraction is cut at its headings and packed into windows of 1,200 characters, the common production default. No chunk carries its heading path: if a figure needs its section title to mean anything, the extraction has to have put that title within reach. |
| Retrieve | BM25 over the chunks, top four for each question, for every extraction alike. |
| Answer | A model answers each question from those four chunks and nothing else, twice: once from each extraction, labelled X and Y in an order randomised per question. It is told to say NOT FOUND rather than guess, and to say NOT FOUND when a figure is there but the chunk does not say which plan, model or case it belongs to. |
| Grade | A second model marks each answer against the gold answer as correct, partial, wrong or not found, without knowing which side is which. Correct counts one, partial a half. |
| Audit | A stronger model re-graded 242 items and checked their gold answers. It agreed with the first grader on 99.0% of grades and corrected 9 gold answers. Later rounds were graded against the corrected gold. |
Why models rather than people? 639 questions, three columns, answered and graded more than once as the extraction improved, is several thousand judgements. A model does them the same way every time, and the audit says how often it is right. The other study on this site uses no model at all, and we keep it that way where the thing measured allows it.
What is the noise in this? Firecrawl's documents were fetched once and never changed, but its answers were regenerated beside every Kerf round. Its score on the re-answered pages moved by up to 1.8 points between rounds with nothing on its side different, so a gap inside that is not a result. The gaps reported below are larger than it.
05a
What came back
Every figure below is written by the harness's exporter from the saved grades, and nothing on this page is typed by hand. The orange dot marks the best of the three where the measure has one.
| Measure | Kerf | Kerf, optimiseForRAG | Firecrawl | What it means |
|---|---|---|---|---|
| Answered correctly | 80.8% | ● 83.7% | 78.0% | Of 639 questions. A partial answer counts a half. |
| Not found | 17.8% | ● 15.0% | 20.8% | The answer was on the page and the retriever never handed it over, or it was not on the page at all. |
| Wrong | 0.9% | ● 0.8% | ● 0.8% | A confident answer that contradicts the page. The figure that matters most and moves least. |
| Pages ahead of Firecrawl | 29 | 32 | – | Of 110 pages, how many this column scored higher on than Firecrawl did. |
| Pages behind Firecrawl | 11 | 8 | – | And how many it scored lower on. The rest were level. |
| Median time a page | ● 1,634ms | – | 2,036ms | Wall clock, as the caller sees it, on the last fetch of every page. The optimiseForRAG column reuses those fetches, so it has no time of its own. |
| Median characters | 18,987 | – | 33,370 | How much Markdown a page became, before any shaping; the shaped output was not measured separately. Less that answers the same questions is fewer tokens to index. |
639 questions over 110 pages. Firecrawl's column is its answers from the same rounds as Kerf's optimiseForRAG column; beside the plain Kerf column it scored 77.5%.
Read the three columns together. Kerf as called answers 80.8% against Firecrawl's 77.5% on the same rounds. The option adds 2.9 points on top, to 83.7%, from the same fetches: nothing new was read, the Markdown was shaped so that a chunk holds whole facts. Wrong answers stay under one per cent for every column, which is the figure to check first in any test that lets a model answer.
05b
By kind of question
Where the difference comes from. The set is built so no one kind of question can carry the result.
| Kind | n | Kerf | Kerf, optimiseForRAG | Firecrawl | What it asks for |
|---|---|---|---|---|---|
| numeric | 181 | 74.3% | 78.2% | 71.0% | A figure: a price, a limit, a version, a count. |
| fact | 139 | 91.0% | 91.0% | 86.7% | A name, a date, a definition, a yes or no. |
| list | 101 | 83.2% | 86.1% | 78.2% | Several items, all of which have to arrive. |
| qualified | 218 | 78.7% | 82.6% | 78.2% | The right answer depends on a qualifier the page sets: which plan, which model, which case. An answer lifted from the wrong part of the page is wrong. |
Correct answers, per kind of question.
The qualified questions are the ones a retrieval pipeline gets wrong most quietly: the figure is on the page, the chunk has a figure, and it is the wrong plan's. They are also where the shaping helps most, because keeping a name beside its figures is what stops a chunk from carrying a number without its qualifier.
06
Two pages in detail
A summary table can hide what a difference actually looks like, so here are two pages, every scored question, with each column's answer as the grader saw it.
github.com/openssl/openssl. A repository page: a README inside a great deal of application chrome, with the counts and the latest release written in by script after the page loads.
| Question | Gold | Kerf | Kerf, optimiseForRAG | Firecrawl |
|---|---|---|---|---|
| What is the latest OpenSSL release version listed on its GitHub repository page? | OpenSSL 4.0.2 | OpenSSL 4.0.2correct | OpenSSL 4.0.2correct | NOT FOUNDnot found |
| What three components make up the OpenSSL toolkit, according to its README? | libssl, libcrypto, and the openssl command-line tool | libssl, libcrypto, and the openssl command line toolcorrect | libssl, libcrypto, and the openssl command line toolcorrect | libssl, libcrypto, and the openssl command line toolcorrect |
| OpenSSL is described as descended from the SSLeay library developed by which two people? | Eric A. Young and Tim J. Hudson | Eric A. Young and Tim J. Hudsoncorrect | Eric A. Young and Tim J. Hudsoncorrect | Eric A. Young and Tim J. Hudsoncorrect |
| Under what license is OpenSSL distributed? | The Apache License 2.0 | Apache License 2.0correct | Apache License 2.0correct | Apache License 2.0correct |
Answers as written by the answering model from the four chunks it was given. Grades as the grader gave them.
apple.com/iphone-18-pro. A product page whose price, colours and specifications are drawn by script into a layout the markup only sketches. The kind of page a direct read gets the words of and not the figures.
| Question | Gold | Kerf | Kerf, optimiseForRAG | Firecrawl |
|---|---|---|---|---|
| What is the starting price of iPhone 18 Pro? | $1199 | From $1199correct | From $1199 (or $49.95/mo. for 24 mo.)correct | From $1199 (or $49.95/mo. for 24 mo.)correct |
| What color options are available for iPhone 18 Pro? | Burgundy, Glacier, Silver, Black | Burgundy, Glacier, Silver, Blackcorrect | Burgundy, Glacier, Silver, Blackcorrect | Burgundy, Glacier, Silver, Blackcorrect |
| How many more hours of video playback does iPhone 18 Pro Max get compared with iPhone 16 Pro Max? | 12 more hours | Up to 12 more hourscorrect | Up to 12 more hourscorrect | NOT FOUNDnot found |
| What are the four aperture settings available on the iPhone 18 Pro's variable-aperture Main camera? | ƒ/1.48, ƒ/1.8, ƒ/2.8, and ƒ/4.0 | NOT FOUNDnot found | ƒ/1.48, ƒ/1.8, ƒ/2.8, ƒ/4.0correct | NOT FOUNDnot found |
| In Apple's carrier deal for iPhone 18 Pro, what is the maximum trade-in credit offered specifically through Verizon (as opposed to AT&T or T-Mobile)? | up to $1020 | Up to $1020correct | Up to $1020correct | Up to $1020correct |
| With faster wired charging, how long does it take iPhone 18 Pro to reach 50 percent battery? | around 15 minutes | Around 15 minutescorrect | Around 15 minutes (with compatible power adapter)correct | Around 15 minutes (with compatible power adapter)correct |
07
Every page, and what each one answered
The whole dataset. Points scored out of the questions asked, for each column; which path Kerf took; how much Markdown came back and how long the fetch took.
| Page | n | Kerf | K +RAG | Firecrawl | K path | K chars | F chars | K ms | F ms |
|---|---|---|---|---|---|---|---|---|---|
| astro.build | 6 | 5/6 | 5/6 | 5/6 | fast | 9,320 | 14,479 | 972 | 2,022 |
| aws.amazon.com/ec2/pricing/on-demand | 6 | 3/6 | 3/6 | 3/6 | fast | 9,919 | 17,150 | 1,795 | 4,547 |
| blog.codinghorror.com/the-best-code-is-no-code-at-al | 6 | 4/6 | 4/6 | 4/6 | fast | 9,352 | 103,824 | 882 | 2,539 |
| blog.regehr.org/archives/213 | 6 | 6/6 | 6/6 | 5/6 | fast | 42,919 | 42,542 | 1,164 | 4,001 |
| books.toscrape.com | 6 | 4/6 | 4/6 | 4/6 | fast | 3,657 | 9,116 | 1,174 | 1,047 |
| bun.sh/docs/installation | 6 | 5/6 | 6/6 | 6/6 | fast | 7,252 | 7,271 | 1,188 | 2,003 |
| caniuse.com/css-grid | 5 | 3/5 | 3/5 | 2/5 | browser | 1,911 | 2,686 | 6,510 | 1,554 |
| creativecommons.org/licenses/by-sa/4.0/legalcode.en | 6 | 6/6 | 6/6 | 6/6 | fast | 18,987 | 22,186 | 1,053 | 1,322 |
| danluu.com/everything-is-broken | 6 | 6/6 | 6/6 | 4/6 | fast | 19,622 | 21,025 | 1,115 | 1,606 |
| datatracker.ietf.org/doc/html/rfc7231 | 6 | 5/6 | 5/6 | 5/6 | fast | 240,387 | 235,220 | 1,206 | 1,935 |
| developer.mozilla.org/en-US/docs/Web/API/Fetch_API | 6 | 6/6 | 6/6 | 6/6 | browser | 7,487 | 21,275 | 9,006 | 1,375 |
| developer.mozilla.org/en-US/docs/Web/CSS/Reference/P | 6 | 5/6 | 6/6 | 4/6 | browser | 12,213 | 15,282 | 11,770 | 1,231 |
| developer.mozilla.org/en-US/docs/Web/HTTP/Reference/ | 6 | 6/6 | 6/6 | 6/6 | browser | 32,905 | 60,405 | 11,659 | 1,714 |
| developer.mozilla.org/en-US/docs/Web/HTTP/Reference/ | 6 | 6/6 | 6/6 | 6/6 | browser | 16,524 | 29,856 | 9,459 | 1,316 |
| doc.rust-lang.org/std/vec/struct.Vec.html | 6 | 3/6 | 5/6 | 3/6 | fast | 227,138 | 578,286 | 2,529 | 3,187 |
| docs.github.com/en/actions/writing-workflows/quickst | 6 | 6/6 | 6/6 | 6/6 | fast | 8,867 | 11,260 | 1,680 | 1,438 |
| docs.python.org/3.12/library/asyncio-task.html | 6 | 6/6 | 6/6 | 6/6 | fast | 38,720 | 64,986 | 899 | 2,474 |
| docs.python.org/3.12/library/pathlib.html | 6 | 6/6 | 6/6 | 6/6 | fast | 45,022 | 75,648 | 1,399 | 1,593 |
| docs.stripe.com/webhooks | 6 | 6/6 | 6/6 | 6/6 | fast | 24,776 | 35,327 | 1,615 | 8,175 |
| en.wikipedia.org/wiki/Comparison_of_file_systems | 6 | 1/6 | 5/6 | 0.5/6 | fast | 213,451 | 261,330 | 1,464 | 4,556 |
| en.wikipedia.org/wiki/Dot-com_bubble | 6 | 6/6 | 5/6 | 5/6 | fast | 60,530 | 181,246 | 890 | 1,763 |
| en.wikipedia.org/wiki/List_of_countries_by_GDP_(nomi | 6 | 4/6 | 4/6 | 5/6 | fast | 56,827 | 170,091 | 1,134 | 2,036 |
| en.wikipedia.org/wiki/List_of_countries_by_populatio | 4 | 4/4 | 3/4 | 4/4 | fast | 71,244 | 150,888 | 1,228 | 1,851 |
| en.wikipedia.org/wiki/PageRank | 6 | 6/6 | 6/6 | 6/6 | fast | 55,200 | 126,602 | 2,166 | 1,908 |
| en.wikipedia.org/wiki/Transmission_Control_Protocol | 6 | 5/6 | 5/6 | 5/6 | fast | 102,316 | 198,018 | 1,637 | 3,480 |
| en.wikipedia.org/wiki/Web_scraping | 6 | 6/6 | 6/6 | 5/6 | fast | 30,113 | 58,021 | 1,415 | 1,532 |
| fastapi.tiangolo.com/tutorial/background-tasks | 6 | 6/6 | 6/6 | 6/6 | fast | 12,051 | 11,124 | 887 | 2,189 |
| gdpr-info.eu/art-17-gdpr | 6 | 6/6 | 6/6 | 6/6 | fast | 2,991 | 3,795 | 1,115 | 3,103 |
| github.com/curl/curl | 5 | 2/5 | 3/5 | 2/5 | browser | 8,638 | 31,082 | 7,649 | 2,807 |
| github.com/microsoft/playwright | 5 | 4/5 | 4/5 | 4/5 | browser | 15,728 | 37,152 | 7,352 | 4,604 |
| github.com/openssl/openssl | 4 | 4/4 | 4/4 | 3/4 | browser | 16,936 | 72,994 | 7,871 | 4,502 |
| github.com/pallets/flask | 5 | 3/5 | 4/5 | 3/5 | browser | 5,712 | 16,228 | 7,693 | 1,808 |
| github.com/postgres/postgres | 6 | 5/6 | 6/6 | 5/6 | browser | 5,787 | 25,091 | 7,693 | 1,844 |
| github.com/pricing | 6 | 5/6 | 6/6 | 2/6 | fast | 24,435 | 39,171 | 1,202 | 1,645 |
| github.com/rust-lang/rust | 3 | 3/3 | 3/3 | 3/3 | browser | 9,114 | 30,785 | 7,234 | 2,033 |
| github.com/sqlite/sqlite | 4 | 4/4 | 4/4 | 4/4 | browser | 24,309 | 41,647 | 9,772 | 2,802 |
| github.com/vercel/next.js | 5 | 4/5 | 3/5 | 3/5 | browser | 14,369 | 184,438 | 7,717 | 3,193 |
| go.dev/blog/loopvar-preview | 6 | 6/6 | 6/6 | 6/6 | fast | 7,754 | 8,696 | 756 | 1,667 |
| html.spec.whatwg.org/multipage/introduction.html | 6 | 6/6 | 6/6 | 6/6 | browser | 54,225 | 60,671 | 8,215 | 1,428 |
| learn.microsoft.com/en-us/azure/virtual-machines/siz | 6 | 6/6 | 6/6 | 5/6 | fast | 45,839 | 82,721 | 1,583 | 2,019 |
| linear.app/pricing | 6 | 6/6 | 6/6 | 6/6 | browser | 2,824 | 3,438 | 9,205 | 2,666 |
| martinfowler.com/articles/microservices.html | 6 | 5/6 | 5/6 | 3/6 | fast | 50,063 | 51,307 | 705 | 1,271 |
| martinfowler.com/bliki/TechnicalDebt.html | 6 | 6/6 | 6/6 | 6/6 | fast | 7,276 | 7,451 | 1,333 | 1,056 |
| motherfuckingwebsite.com | 6 | 6/6 | 6/6 | 6/6 | fast | 3,382 | 3,312 | 913 | 1,343 |
| nextjs.org/docs/app/getting-started/installation | 6 | 6/6 | 6/6 | 6/6 | fast | 18,590 | 14,930 | 1,112 | 2,273 |
| nodejs.org/docs/latest-v22.x/api/fs.html | 6 | 6/6 | 6/6 | 4/6 | fast | 194,860 | 437,934 | 2,858 | 4,114 |
| openai.com/api/pricing | 5 | 0/5 | 0/5 | 0/5 | fast | 15,606 | 17,745 | 4,063 | 3,068 |
| opensource.org/license/mit | 6 | 4.5/6 | 4.5/6 | 2.5/6 | browser | 1,322 | 3,526 | 7,171 | 1,779 |
| overreacted.io/a-complete-guide-to-useeffect | 6 | 5/6 | 5/6 | 5/6 | fast | 65,532 | 71,703 | 1,373 | 2,739 |
| pkg.go.dev/net/http | 6 | 4/6 | 4/6 | 3/6 | fast | 164,012 | 207,354 | 1,983 | 1,792 |
| playwright.dev/docs/locators | 6 | 6/6 | 6/6 | 6/6 | fast | 24,392 | 33,865 | 990 | 2,942 |
| policies.google.com/terms | 6 | 3/6 | 3/6 | 3/6 | fast | 30,155 | 39,296 | 1,202 | 1,780 |
| react.dev/learn/thinking-in-react | 6 | 5/6 | 5/6 | 4/6 | fast | 21,763 | 24,241 | 852 | 1,511 |
| semver.org | 6 | 6/6 | 6/6 | 6/6 | fast | 17,734 | 17,794 | 831 | 1,137 |
| slack.com/pricing | 6 | 5/6 | 5/6 | 5/6 | fast | 23,043 | 36,734 | 3,028 | 2,018 |
| store.google.com/us/product/pixel_10_pro | 6 | 3/6 | 4/6 | 3/6 | fast | 34,931 | 84,311 | 3,937 | 7,009 |
| stripe.com/gb/pricing | 6 | 6/6 | 6/6 | 6/6 | fast | 24,858 | 34,428 | 1,232 | 2,755 |
| supabase.com/pricing | 6 | 2.5/6 | 4/6 | 2.5/6 | fast | 29,149 | 22,468 | 867 | 1,374 |
| tailwindcss.com | 6 | 4/6 | 4/6 | 4/6 | fast | 8,402 | 11,424 | 1,111 | 3,936 |
| unicode.org/emoji/charts/full-emoji-list.html | 6 | 3/6 | 5/6 | 4/6 | fast | 153,581 | 453,177 | 9,116 | 5,890 |
| vercel.com/pricing | 6 | 4/6 | 5/6 | 5/6 | fast | 13,327 | 15,609 | 1,501 | 4,185 |
| vite.dev/guide | 6 | 5/6 | 5/6 | 5/6 | fast | 9,281 | 13,149 | 1,288 | 1,850 |
| www.adobe.com/creativecloud/plans.html | 5 | 4/5 | 4/5 | 3/5 | browser | 3,911 | 23,890 | 10,441 | 726 |
| www.anthropic.com/pricing | 6 | 2/6 | 3/6 | 3/6 | fast | 17,576 | 24,317 | 2,934 | 1,958 |
| www.apache.org/licenses/LICENSE-2.0 | 6 | 5.5/6 | 5.5/6 | 6/6 | fast | 11,156 | 11,514 | 688 | 4,522 |
| www.apple.com/airpods-pro | 6 | 5/6 | 5/6 | 6/6 | browser | 16,949 | 37,787 | 20,748 | 1,838 |
| www.apple.com/iphone-17 | 6 | 4/6 | 4/6 | 4/6 | browser | 27,079 | 52,999 | 13,314 | 2,442 |
| www.apple.com/iphone-18-pro | 6 | 5/6 | 6/6 | 4/6 | browser | 117,608 | 64,751 | 9,387 | 2,343 |
| www.apple.com/iphone/compare | 6 | 1/6 | 1/6 | 1/6 | fast | 47,761 | 299,419 | 2,667 | 2,519 |
| www.apple.com/macbook-pro | 6 | 6/6 | 6/6 | 4/6 | browser | 15,488 | 46,521 | 9,736 | 2,429 |
| www.apple.com/watch | 6 | 2/6 | 1.5/6 | 5/6 | browser | 5,013 | 5,483 | 9,433 | 1,869 |
| www.atlassian.com/software/confluence/pricing | 6 | 5/6 | 5/6 | 5/6 | browser | 14,797 | 31,322 | 13,733 | 2,964 |
| www.atlassian.com/software/jira/pricing | 6 | 4/6 | 4/6 | 4/6 | browser | 13,368 | 37,918 | 12,226 | 695 |
| www.canva.com/pricing | 6 | 5/6 | 5/6 | 5/6 | browser | 16,176 | 24,753 | 12,378 | 772 |
| www.cloudflare.com/plans | 6 | 4/6 | 5/6 | 4/6 | fast | 12,501 | 26,950 | 1,268 | 1,956 |
| www.dell.com/en-us/shop/dell-laptops/scr/laptops | 6 | 5.5/6 | 5.5/6 | 6/6 | fast | 20,182 | 66,570 | 1,631 | 7,893 |
| www.digitalocean.com/pricing/droplets | 6 | 5/6 | 5/6 | 5/6 | fast | 96,451 | 26,982 | 781 | 1,558 |
| www.dropbox.com/plans | 6 | 3/6 | 3/6 | 3/6 | fast | 6,250 | 8,786 | 1,634 | 3,898 |
| www.figma.com/pricing | 6 | 3/6 | 5/6 | 3/6 | fast | 29,823 | 37,178 | 1,142 | 3,544 |
| www.garmin.com/en-US/p/886785 | 6 | 6/6 | 6/6 | 6/6 | browser | 1,477 | 6,085 | 19,533 | 2,097 |
| www.gnu.org/licenses/gpl-3.0.en.html | 6 | 6/6 | 6/6 | 6/6 | fast | 35,583 | 39,040 | 842 | 1,010 |
| www.hetzner.com/cloud | 6 | 4/6 | 5/6 | 3/6 | browser | 8,857 | 32,184 | 9,991 | 3,347 |
| www.hubspot.com/pricing/marketing | 6 | 2/6 | 4/6 | 5/6 | browser | 16,884 | 17,732 | 16,170 | 794 |
| www.iana.org/assignments/http-status-codes | 6 | 6/6 | 6/6 | 6/6 | fast | 4,482 | 7,801 | 1,218 | 2,136 |
| www.ikea.com/us/en/p/billy-bookcase-white-00263850 | 5 | 5/5 | 5/5 | 5/5 | browser | 37,380 | 55,916 | 13,526 | 2,745 |
| www.indeed.com/career-advice | 5 | 0/5 | 0/5 | 4/5 | – | 0 | 33,370 | 8,284 | 3,591 |
| www.joelonsoftware.com/2000/04/06/things-you-should- | 6 | 4/6 | 4/6 | 4/6 | fast | 8,969 | 9,329 | 937 | 1,687 |
| www.kalzumeus.com/2012/01/23/salary-negotiation | 6 | 6/6 | 6/6 | 6/6 | browser | 41,653 | 43,003 | 6,663 | 2,314 |
| www.lg.com/us/tvs/lg-oled65c5pua-oled-4k-tv | 6 | 4.5/6 | 4/6 | 3/6 | browser | 32,012 | 51,021 | 25,536 | 5,086 |
| www.microsoft.com/en-us/surface | 6 | 5/6 | 5/6 | 5/6 | fast | 11,360 | 14,098 | 1,074 | 6,038 |
| www.mongodb.com/pricing | 6 | 5/6 | 4/6 | 4/6 | fast | 27,041 | 22,485 | 2,960 | 4,907 |
| www.netcup.com/en/server/vps | 6 | 6/6 | 6/6 | 5/6 | fast | 14,746 | 19,613 | 2,570 | 4,357 |
| www.notion.com/pricing | 6 | 5/6 | 5/6 | 5/6 | fast | 21,743 | 23,414 | 893 | 685 |
| www.paulgraham.com/ds.html | 6 | 6/6 | 6/6 | 6/6 | fast | 25,240 | 28,325 | 1,082 | 1,708 |
| www.paulgraham.com/makersschedule.html | 6 | 6/6 | 6/6 | 6/6 | fast | 6,859 | 8,794 | 1,573 | 2,210 |
| www.postgresql.org/docs/16/datatype-json.html | 6 | 6/6 | 6/6 | 6/6 | fast | 30,374 | 35,973 | 1,447 | 4,047 |
| www.postgresql.org/docs/16/sql-select.html | 6 | 6/6 | 6/6 | 6/6 | fast | 63,209 | 71,366 | 1,037 | 2,447 |
| www.rei.com/product/222507/rei-co-op-flash-55-pack-m | 4 | 3/4 | 3/4 | 2/4 | browser | 4,756 | 20,112 | 36,304 | 6,061 |
| www.rfc-editor.org/rfc/rfc1925.html | 6 | 6/6 | 6/6 | 6/6 | fast | 4,327 | 4,280 | 909 | 6,178 |
| www.rfc-editor.org/rfc/rfc6749.html | 6 | 5/6 | 5/6 | 5/6 | fast | 163,533 | 164,070 | 982 | 1,124 |
| www.rfc-editor.org/rfc/rfc8259.html | 6 | 6/6 | 6/6 | 6/6 | fast | 28,359 | 28,441 | 1,402 | 1,039 |
| www.rfc-editor.org/rfc/rfc9110.html | 6 | 5/6 | 4/6 | 4/6 | fast | 442,756 | 750,991 | 2,092 | 3,852 |
| www.rfc-editor.org/rfc/rfc9309.html | 6 | 6/6 | 5.5/6 | 5.5/6 | fast | 22,753 | 34,065 | 1,680 | 1,322 |
| www.samsung.com/us/smartphones/galaxy-s25-ultra | 6 | 6/6 | 6/6 | 6/6 | browser | 42,762 | 90,436 | 13,918 | 4,331 |
| www.shopify.com/pricing | 6 | 5/6 | 4.5/6 | 4.5/6 | fast | 26,797 | 50,274 | 1,813 | 1,241 |
| www.spotify.com/us/premium | 5 | 5/5 | 5/5 | 4/5 | fast | 6,538 | 6,592 | 815 | 1,625 |
| www.tesla.com/modely | 6 | 2/6 | 2/6 | 2/6 | browser | 19,827 | 50,996 | 26,637 | 5,144 |
| www.twilio.com/en-us/sms/pricing/us | 6 | 6/6 | 6/6 | 6/6 | fast | 10,541 | 13,746 | 948 | 3,821 |
| www.w3.org/TR/WCAG22 | 6 | 6/6 | 6/6 | 6/6 | fast | 133,606 | 285,907 | 1,247 | 1,817 |
| zoom.us/pricing | 6 | 4/6 | 4/6 | 3/6 | browser | 14,092 | 202,020 | 20,146 | 879 |
Points are correct answers plus half for each partial answer, out of the questions scored on that page. Characters and time are from the last fetch of each page.
08
What this does not show
The choices that shape the result, kept in one place so the figures above can stand without qualifying themselves.
01
Write
Questions are written from both extractions of a page, blind to which is which, and frozen with their gold answers before anything is answered.
02
Answer and grade
Chunks are retrieved the same way for every extraction, answered blind, graded blind against the gold. Nothing is scored by hand.
03
Round
When Kerf changes, only the pages whose extraction changed are answered and graded again, both columns together, so a page is never scored against a stale rival.
Models answer and grade. A person did not read 639 answers three times over. The audit above says how far to trust that, and the noise figure says how much of a gap it takes to mean anything. Both are published for exactly that reason.
The rounds are not one fetch. Kerf was fetched again as it improved, over 10 rounds, and each page's answers come from the latest round in which its extraction changed. Firecrawl was fetched once and its documents never changed; its answers were regenerated beside each Kerf round so the two columns on any page always come from the same sitting. This dataset was also used during Kerf development, and as such, treat it as a transparent development benchmark rather than a held-out generalisation test. We will run the final implementation unchanged against a fresh frozen set.
One chunker, one retriever. 1,200-character chunks cut at headings, and BM25 over them, are a common production default and not the only one. We checked the shaping at 500 and 2,000 characters too and the direction held, but a pipeline with a different chunker or an embedding retriever will see a different size of gap.
Kerf was called with links off. That is what a retrieval caller sends, since addresses match nothing anyone asks, and it is the setting the option is meant to sit beside. Firecrawl was called at its defaults, which keep links.
The shaping trades structure for retrieval. Fewer headings and no Markdown tables come back with the option on. On the structure measures our other study scores, the same Markdown does worse, which is why the option is off by default and this page shows both columns.
09
How it was run
Four steps, in this order, with everything saved between them so the scoring can be re-run as often as the argument needs.
| 1 | Fetch every page with each product from the United States, and save what came back verbatim. |
| 2 | Write up to six questions per page from the two extractions, blind, and freeze them with their gold answers. |
| 3 | Chunk, retrieve, answer and grade, blind, both columns in the same sitting. Audit the first round. |
| 4 | Export every grade to the dataset behind this page. Nothing on it is typed. |
Firecrawl fetched 13 September 2026, Kerf fetched 14 September 2026, scored 14 September 2026. Firecrawl is called through its own public API at its own defaults, with the location set to match ours. Naming another product is not a claim about anything except what it returned on these pages on that day.