A PDF is a document that Google downloads, opens, reads and indexes like any other page: in the results it appears with its own URL, its own title and its own snippet. SEO for PDFs is the set of choices that decide whether that file reaches the first page, stays invisible or is better taken out of the index.
Guides on the subject have been repeating the same list of tips for years: hyphens in the file name, metadata filled in, selectable text. Few of them say how often PDFs actually rank, and fewer check what ChatGPT does with those files. Below you will find both measurements, taken on 26 September 2026.
How many PDFs are really on the first page in the UK
We pulled the first pages of six UK queries with DataForSEO (Google UK, desktop, 26 September 2026) and counted how many organic results were files with a .pdf extension. Of 52 results, 22 were PDFs: 42.3%.
| Query (UK, desktop, 26 September 2026) | Organic results | Of which .pdf files | Position of the first PDF |
|---|---|---|---|
| heat pump technical data sheet | 9 | 9 | 1 |
| forklift operator manual | 9 | 2 | 1 |
| pdf seo | 9 | 7 | 1 |
| how to optimise pdf for seo | 7 | 1 | 6 |
| declaration of performance ukca marking | 9 | 3 | 2 |
| business grants 2026 | 9 | 0 | none |
| Total | 52 | 22 (42.3%) |
The clearest case is "heat pump technical data sheet": all nine organic results are PDFs, from Regulus, Wondrwall, General, FETA, Daikin, Solplanet, Kotlarnica, Dimplex and Rinnai UK. There is no HTML page on the first page at all. "pdf seo" is a curious one: Google reads it as a request for SEO guides in PDF form, and seven of the nine results are exactly that. At the other extreme, "business grants 2026" has not a single PDF on the first page: the answers come from data.gov.uk, a regional growth hub, accountants and funding advisers who rewrite those grants as web pages.
The difference is not the format. It is who has done the work. Where someone has written a page that answers the question, the page wins. Where the only document containing the rating-plate data is the manufacturer's file, the file ranks. It is the same pattern we see on industrial projects: in the Macropix case, none of the nine keywords that bring in the most enquiries contains the company name, and the biggest one ("ledwall") is worth 4,400 searches a month in Italy. If you do not get there with a page, someone else's PDF will.
What Google does with a PDF
Google's documentation on indexable file types lists PDF among the supported encoded formats and says that "Google can index the content of most text-based files and certain encoded document formats". So a PDF is not penalised from the outset: it is treated as a document with its own text content.
The differences from an HTML page lie elsewhere, and there are four of them.
- The text has to exist as text. A scan saved as a PDF is an image inside a container: without an OCR pass there is nothing to index.
- A PDF cannot carry structured data. The formats Google documents for structured data are JSON-LD, Microdata and RDFa, and they all live in HTML. A file that ranks will never get a rich result, and no schema markup will help it.
- The title shown in the SERP comes from the document properties, not from a tag. If the Title field is empty, Google takes whatever it finds: often the export file name, something like "Microsoft Word - draft_final2".
- Links inside a PDF work as links. The GDPR PDF attracts about 77,000 links from 823 referring domains and the PDF of Google's SEO starter guide 3,370 links from 754 domains, according to Ahrefs' analysis (Ahrefs index, global data).
Those two numbers explain why a citable PDF is an asset and not a problem: a document that others link to brings authority to the domain even though it is not a page. The rest of the crawling rules apply unchanged, including the sitemap, where PDF URLs can be listed like any other URL.
ChatGPT opens PDFs and quotes their numbers
On 26 September 2026 we asked ChatGPT, with web search on, in English and with a UK location, a question straight out of a technical manual: what is the COP of a Daikin Altherma 3 M 8 kW heat pump.
The answer gave a COP of about 4.6 at 7 °C outdoor air with 35 °C flow water, about 3.5 at 55 °C flow, and seasonal values (SCOP) of 4.5-4.6 at 35 °C and about 3.3 at 55 °C. It cited four sources: the Heat Pump Keymark certification database, a Danish retailer's product page, an article on daikin.co.uk and a PDF, the Daikin Altherma 3 technical specifications flyer on daikin.co.uk. The PDF was the source for the 55 °C figure and for both seasonal values. We had run the same question in Italian with an Italian location the same day: there, two of the three sources cited were PDFs on daikin.it, including a manufacturer's declaration signed and dated 9 April 2026.
Not pages that talk about a PDF: the files themselves, opened and read. OpenAI's crawler documentation lists four user agents (OAI-SearchBot, OAI-AdsBot, GPTBot and ChatGPT-User) and does not state which file formats are downloaded and processed, so the only way to know is to measure it query by query, as we did here and as we describe in the guide on how to do SEO for ChatGPT.
The practical consequence is awkward for anyone who has always said to put everything in HTML: in a technical sector, the table inside the product datasheet is exactly the part that gets read and repeated. If that table is a scanned image, the model does not see it and cites the competitor who exported the text.
The six things that decide whether a PDF ranks
In order of impact, from the heaviest to the lightest.
1. Selectable text, not an image
Open the file and try to select a line with the mouse. If nothing happens, to Google that document is empty. It happens to every catalogue laid out as graphics and exported as an image, and to every scan of a paper document. The fix is to export from the source file instead of the scanner, or to run OCR (Acrobat, ocrmypdf, Tesseract).
2. The Title field in the document properties
This is the PDF's title tag. In Acrobat it is under File, Properties, Description; in LibreOffice and Word under the document properties. Fill in Title and Subject with the phrase a person would search for, not the internal product code. "Air-to-water heat pump 8 kW, datasheet and COP curves" works. "DEP-0043 rev.4" does not.
3. The file name and the folder it sits in
A PDF's URL is its file name, and it stays visible in the snippet. Words separated by hyphens, no spaces encoded as %20, and no revision dates in the name if you plan to update it: every rename is a new URL that loses everything it had built up. Keep them in a dedicated folder, such as /documents/, so they can be isolated with a single rule when needed.
4. The heading structure inside the document
The Heading 1, Heading 2 and Heading 3 styles in your word processor become structural tags in the exported PDF. A document formatted by hand with bold and 18-point type is, structurally, a single block of text. The difference shows when Google has to work out what page 14 of 60 is about.
5. Alt text on images and file size
Images inside a PDF accept alternative text exactly like those on a page, and the same rules apply that we use for alt text and image optimisation. On size, the yardstick is the user: nobody opens a 40 MB catalogue downloaded over mobile data at a trade show, and download time is part of the page experience like any other file served from your domain.
6. Inbound and outbound links
An isolated PDF, reachable only from a "Download" button at the bottom of a page, stays on the edges of the site. Treat it as a real resource: link to it from the text of relevant pages using the same logic as internal links, and put links inside it to the pages on the site that complete the document. Someone opening a datasheet downloaded two years ago should be able to get back to the current configurator.
When to take a PDF out of the index
Three recurring situations: the PDF that repeats an existing page word for word and takes its clicks, the old price list that keeps appearing instead of the new one, and forms that pick up informational searches and send people to a sheet to fill in.
robots.txt fixes none of the three: it blocks crawling, not indexing, and a blocked URL can stay in the SERP without a snippet. For a non-HTML file, noindex is declared in the response header, with X-Robots-Tag. Google documents the two rules like this:
# Apache
<Files ~ "\.pdf$">
Header set X-Robots-Tag "noindex, nofollow"
</Files>
# NGINX
location ~* \.pdf$ {
add_header X-Robots-Tag "noindex, nofollow";
}
If instead the PDF must stay downloadable but should not compete with the HTML page that says the same things, the route is a canonical declared in the HTTP header, which Google describes in its guide on consolidating duplicate URLs. The syntax follows RFC 5988 and requires an absolute URL:
Link: <https://www.example.co.uk/documents/white-paper.pdf>; rel="canonical"
Bear in mind that the canonical remains a hint and not a command, with PDFs as with pages: the logic is the same as described in the guide on the canonical URL.
PDF or HTML page: how to decide
The criterion we use on B2B projects is simple: the PDF is the format for people who need to take the document away, the page is the format for people who need to find it. A specification the designer attaches to the tender, a datasheet the engineer prints and takes to site, a set of accounts someone files away: these are files, and they should be published as files. A guide, a product comparison, an answer to a question: these are pages.
When you need both, the set-up that holds is an HTML page containing the data as text with the PDF linked at the bottom, carrying the same version of the numbers. It costs an extra half hour per document and solves duplication, the structured data problem and updating in advance: when the price list changes, you change the page and replace the file at the same URL. It is the kind of decision that comes out of an SEO audit on an industrial site, where half of the useful content is often locked in a folder of attachments.
An unreadable PDF is also an accessibility compliance problem
In the UK, the Public Sector Bodies (Websites and Mobile Applications) (No. 2) Accessibility Regulations 2018 require public sector sites to meet WCAG 2.2 AA, and PDFs are covered: the only exemption is for documents published before 23 September 2018, and even those must be made accessible if users need them to use a service. For private businesses, the same government guidance notes that all UK service providers have a legal obligation to make reasonable adjustments under the Equality Act 2010. If you sell to consumers in the EU, the European Accessibility Act (Directive 2019/882) has also applied since 28 June 2025.
The interesting part is that the fixes are the same. Real text instead of a scan, correct reading order, heading tags, alternative text on figures, document language declared: it is the same list that makes a file readable by Googlebot and citable by a language model. Anyone who has already fixed their PDFs for accessibility has done almost all of the SEO work, and vice versa. We have gathered the obligations and who they apply to in the guide on web accessibility.
How to check whether a PDF is bringing in anything
The quick check is the site operator combined with the file type: site:yourdomain.co.uk filetype:pdf shows how many files the search engine holds in its index. The number is approximate, but if 400 come up when you have published 30, you have a problem with duplicates generated by the CMS.
The real data is in Google Search Console: PDF URLs appear in the Performance report like any other page, with clicks, impressions and queries. Filter pages by "pdf" and look at two things: which files get impressions on queries one of your pages should be capturing (there the PDF is cannibalising you) and which get clicks on queries none of your pages covers (there the PDF is telling you which page to write). In GA4 a PDF does not generate a page view, so the download has to be tracked as an event on the page that hosts it, otherwise the file disappears from the reports.
On an industrial site this exercise takes an afternoon and usually produces a list of five or six pages to write. If you would rather have someone else do it, it is part of what we do in our SEO consultancy projects, and in the case studies we publish you will find the starting and finishing numbers.
Frequently asked questions
Something the industry rarely says
A PDF that ranks well is almost always the symptom of a page nobody has written. That goes for you and for your competitors: if the first page for a commercial query is full of manufacturers' datasheets, that query is unguarded, and whoever arrives with a well-made page takes it. It is the opposite of the usual reading, which looks at your own PDF and asks how to optimise it. The more profitable question is the other one: which PDFs, including other people's, are occupying a place that should belong to a page of yours.