Short answer
A PDF (Portable Document Format) is not an image. It is a small database of objects describing a set of pages, with the text, the vector shapes, the images and usually the fonts themselves packed inside one file. Because the fonts travel with it, a PDF looks the same on every device, which a word processor document does not. The cost of that fidelity is that a PDF is fixed: it does not reflow to fit a phone, and it is awkward to edit.
Everybody uses PDFs and almost nobody knows what one is. The useful mental model is not “a picture of a document” and not “a locked Word file”. A PDF is a small database of objects that together describe a set of pages, with the fonts, the images and the vector shapes it needs packed inside the same file, so that the page can be drawn identically anywhere.
That one design decision explains nearly everything people find strange about PDFs: why they look right on every machine, why they will not reflow on a phone, why editing one is awkward, why some are searchable and some are not, and why a three page document can be forty megabytes.
What is actually inside a PDF
Open a PDF in a text editor and the first line is %PDF-1.7 or similar. After that, most of it is unreadable, but the structure is simple enough to describe.
A PDF is a list of numbered objects. Some objects are dictionaries of key and value pairs, some are streams of compressed bytes. There is a catalogue object that points to a page tree, the page tree points at the individual pages, and each page points at two things: a content stream, which is the list of drawing instructions for that page, and a resources dictionary naming everything the instructions refer to.
The content stream itself is small and almost readable. It says things like “set this font at this size”, “move to this position”, “show this string”, “fill this rectangle”, “draw the image called Im1 into this box”. The coordinates are in points, seventy-two to the inch, counted from the bottom left corner of the page, which is why PDF measurements are always in inches rather than pixels.
At the end of the file is a cross-reference table: a list saying, for every object, how many bytes into the file it starts. That is how a reader opens a 900 page document instantly without reading it all: it reads the table, then jumps straight to the object it wants.
Why it looks the same everywhere
A word processor document is a set of instructions for laying out text: this paragraph, in this font, at this size. The computer opening it has to find the font, measure every character and decide where the lines break. If the font is missing, a substitute is used, the measurements change, and the document moves. Everyone has seen a CV arrive with its last line pushed onto a second page.
A PDF has already done that work. The line breaks are decided, each run of text has a position, and the font itself is normally embedded in the file: not a reference to Helvetica, but the actual outlines of the characters used. A reader on a phone in 2035 will draw the page the way its author saw it, because everything the page needs is inside it.
That is the whole promise of the format, and it is why PDF is what a printer, a government form, a bank statement, a contract and a plane ticket all use.
The cost is the same thing said from the other side. Because every decision is baked in, a PDF cannot reflow to fit a small screen, which is why reading one on a phone means pinching and scrolling. And because the text is positioned character by character rather than stored as paragraphs, editing a sentence in the middle of a page is genuinely hard: the editor has to re-flow something that was never stored as flowing text.
Why some PDFs are searchable and some are not
This is the single most useful thing to understand about PDFs, because it explains most of the frustration.
There are two completely different things people call a PDF.
A born-digital PDF was made by a program: Word, InDesign, a browser’s print dialogue, an accounting system. Its pages contain real text objects, so you can select a sentence, copy it, search for a word, and the file is usually small, because text and vector shapes cost almost nothing to store.
A scanned PDF was made by a scanner, a phone camera or a photocopier. Its pages contain one big image each, and nothing else. The letters on the page are shapes made of pixels, exactly as they are in a photograph of a page. You cannot select them, you cannot search them, and the file is large because it is a set of photographs.
Making a scanned PDF searchable needs OCR, optical character recognition, which looks at the shapes and guesses the characters. Good OCR is very good, and it is still guessing: it writes an invisible layer of text on top of the picture, so you can select and search, but what you select is the machine’s reading of the page rather than the page itself.
The quick way to tell which you have: try to select a word. If the cursor draws a text selection, it is born-digital or has an OCR layer. If it draws a rectangle, you have a picture.
Images inside a PDF
A PDF can hold images in several ways, and the one it picks matters if you care about quality or size.
The most common is DCTDecode, which is JPEG. The image data inside the PDF is exactly what would be in a .jpg file: the same compressed bytes, in the same arrangement. This has a pleasant consequence that most conversion tools ignore. Putting a JPEG photo into a PDF should be a copy rather than a conversion, because the target format already speaks JPEG. Tools that decode the picture and re-encode it, which is the easy way to write one, spend a generation of JPEG quality for nothing. Our JPG to PDF page copies instead.
The second is FlateDecode, which is the same deflate compression a PNG uses, usually with the same row-by-row prediction. This is lossless: every pixel is stored exactly. It is the right choice for a screenshot, a logo or a diagram, and the wrong choice for a photograph, where it produces an enormous file.
There are others, including CCITT fax encoding for black and white scans and JPEG 2000 for archives, but those two cover almost everything you will meet.
A PDF image can also carry a soft mask: a second, greyscale image saying how opaque each pixel is. That is the same idea as the alpha channel in a PNG, and it is how transparency works in a PDF. Most readers draw white paper behind the page, so a transparent PDF looks ordinary until you place it in a design tool, which is the moment it matters. What a transparent background actually is explains the idea in general.
Why a PDF can be enormous
Almost always for one of three reasons.
It is a scan. Every page is a photograph. A 20 page scan at 300 dots per inch is 20 pictures of about 2,480 by 3,508 pixels, and if the scanner saved them losslessly rather than as JPEG, that is easily 100 MB.
Someone put photos in at full size. A phone photo is 12 megapixels; ten of them in a document is a large file however it is stored, unless the tool reduced them, which is its own problem because then you cannot get the detail back.
Fonts. Embedding a full font family adds a few hundred kilobytes, and embedding several adds a few megabytes. This is why a two page letter can be 5 MB.
Shrinking a PDF, therefore, mostly means shrinking the pictures inside it: re-compressing them as JPEG, or reducing their resolution. Both are one-way. It is worth knowing which of the three reasons applies to your file before reaching for a compressor, because if the size is fonts, no image compressor will help.
Turning a PDF into pictures, and back
Because a page is a program rather than a picture, turning a PDF into images means running that program at a resolution you choose. That is why every PDF-to-image tool asks how big you want the result, and why asking in dots per inch is the honest way to ask: an A4 page is 8.27 by 11.69 inches, so 150 dpi gives 1,240 by 1,753 pixels and 300 dpi gives 2,480 by 3,508. What dpi actually means goes into the arithmetic, which turns out to matter for scanning too.
Two other jobs come up constantly and are neither of those. Putting several documents together is a copy of objects from one file into another, so nothing is redrawn and nothing is lost: how to combine PDF files covers the built-in way on each device and the things a merge quietly breaks. And making one smaller means shrinking the pictures inside it, which how to compress a PDF goes through, including how to tell in advance whether yours will shrink at all.
Going the other way, pictures into a PDF, is simpler: each picture becomes a page, and the only real decisions are the page size, how the picture sits on it, and whether the image data is copied or re-compressed. Does converting to PDF lose quality is about that last one, and the answer depends entirely on which tool you use.
A short history, and why it matters now
Adobe invented PDF in 1993, as part of a project explicitly about making documents that looked the same everywhere. For its first fifteen years it was Adobe’s format, read with Adobe’s free reader.
In 2008 Adobe published the specification as an open ISO standard, ISO 32000. That is the reason every browser can now open a PDF without a plugin: the format belongs to nobody, anyone may implement it, and Mozilla wrote a renderer in JavaScript, pdf.js, which is what Firefox uses and what this site uses on its own PDF pages.
The practical consequence for you is that PDF is a safe format to keep things in. It is not owned by a company that might stop supporting it, it is fully documented, and there are several independent implementations. Of the formats your files are likely to be in, it is one of the least likely to become unreadable.
Is a PDF safe?
Mostly, with one thing worth knowing. The format allows a PDF to contain JavaScript and to embed other files, and both have been used to carry malware. The risk today is much smaller than it was, because modern readers run PDFs in a sandbox and have scripting off by default, and because the browser you already have is one of those readers.
The sensible habits: open unexpected attachments in your browser rather than in a desktop reader, keep the reader updated, and be suspicious of a PDF that asks you to enable anything.
There is a second kind of safety, which is about where your document goes. A PDF is far more likely than a photograph to be something you would not want copied: a contract, a payslip, a bank statement, a medical letter, a passport scan. People paste those into free online converters without a thought, and those services upload the file to a machine they control and keep it for a period described in a policy nobody reads. Is it safe to upload photos online walks through how to judge that, and the answer applies doubly to documents.
The free way
Turn pictures into a PDF, or a PDF into pictures, on your own device
Both directions run in your browser on getPNG: nothing is uploaded, there is no page limit and no account. Photos are copied into the document rather than re-compressed, transparency is kept if you want it, and going the other way you choose the resolution in dots per inch and which pages you need.
Pictures and PDFs: every guide
What a PDF actually is, how to get pictures into one and out again on any device, and the two things that decide whether the result is any good: whether the tool re-compressed your photo, and what resolution you asked for.
- How to turn a picture into a PDF, on any deviceThe built-in way on iPhone, Android, Windows and Mac, and where each runs out
- Does converting an image to PDF lose quality?Why most converters re-compress a photo that could have been copied
- DPI explained: what resolution to scan, print and export atWhat to scan, print and export at, and why the number in the file means nothing
- How to get the images out of a PDFRendering a page and extracting a placed picture are two different jobs
- How to combine PDF files into one, on any deviceThe built-in way on each device, getting the order right, and what a merge breaks
- How to compress a PDF, and why some will not shrinkWhat actually makes a PDF big, and how to tell whether yours will shrink at all
The tools for this: Any pictures into one PDF, JPG to PDF, without re-compressing, PNG to PDF, losslessly, PDF pages out as PNG, PDF pages out as JPG, Merge PDFs into one, Split a PDF up, Make a PDF smaller.
The short version
A PDF is a page description, not a picture: objects, drawing instructions, embedded fonts and embedded images, wrapped in a file with an index at the end. That design is why it looks identical everywhere, why it will not reflow, why some are searchable and some are not, and why the pictures inside it are the thing that decides its size.
Once you know whether yours is born-digital or a scan, and whether its bulk is images or fonts, almost every PDF question answers itself.
Questions
Is a PDF an image?
No. A page in a PDF is a list of drawing instructions: put this text in this font at this position, fill this shape, draw this image here. Some PDFs contain nothing but a photograph of a page, which is what a scanner makes, and those behave like images. The difference is why text is selectable in some PDFs and not others.
Why does a PDF look the same on every computer?
Because the fonts are usually embedded in the file, along with exact positions for every character. A Word document asks the computer to find a font and lay the text out again, so it moves when the font is missing. A PDF has already made every decision and carries what it needs.
Why can I not select the text in some PDFs?
Because there is no text in them. A scanned or photographed document is a picture of a page, so the letters are shapes, not characters. Making them searchable needs OCR, which reads the shapes and guesses the words.
Is a PDF smaller than the images inside it?
Usually about the same. A PDF stores photographs as JPEG in the same form a .jpg file does, so putting photos into a PDF adds the page structure and little else, unless the tool re-compresses them on the way in.
Can a PDF have a transparent background?
Yes. The format has supported transparency since PDF 1.4, and an image inside one can carry a soft mask, which is the same idea as a PNG's alpha channel. Most readers draw white paper behind a page anyway, so you usually only see it when the file is placed in a design tool.
Is a PDF safe to open?
The format allows JavaScript and embedded files, which is why attachments from strangers deserve care. Modern readers, including the one built into every browser, run PDFs in a sandbox with scripting off by default, which removes most of the risk.
Who owns the PDF format?
Nobody, any more. Adobe created it in 1993 and published it as an open ISO standard in 2008 (ISO 32000). Anyone may read or write PDFs without permission or a licence fee, which is why every browser can now show one.