getPNGgetSVG
Uses9 min read

What AI upscaling can and cannot do

How a super-resolution model actually works, why it draws plausible detail rather than recovering real detail, and where that difference matters.

The getPNG team23 Sept 2026

Short answer

AI upscaling works: a super-resolution model trained on millions of pairs of large and small photographs has learned what skin, brick, fabric and hair look like at the larger size, and draws that instead of smearing. What it cannot do is recover information that was never recorded. It produces a plausible face rather than the real one, and plausible characters on a number plate rather than the right ones. That is fine for a poster and wrong for anything that will be relied on as evidence.

“Enhance” has been a joke in television scripts for thirty years, and then it partly came true. Modern super-resolution models produce results that would have been impossible in 2015. They also do something quite specific, and understanding exactly what it is decides whether you should use one for a particular job.

How a super-resolution model is trained

Take a large, sharp photograph. Shrink it. Now you have a pair: a small picture and the large one it came from. Do that a few million times over a varied collection of photographs and you have a training set.

The network is then shown the small picture and asked to produce the large one. Whatever it produces is compared with the real large version, the difference is measured, and the network’s weights are nudged to reduce it. Repeat for days.

What the network ends up learning is a very rich statistical model of how detail behaves. It learns that a smooth grey gradient with a certain kind of fine noise is skin, and what pores look like on skin. It learns that a repeating brown pattern with a particular structure is brick, and what mortar looks like between bricks. It learns that a group of thin dark streaks against a bright background is hair, and how individual strands separate at this scale.

At the end of training, it is a machine for answering one question: given this small patch, what is the most likely large patch it came from?

That question has more than one answer

Here is the part that matters, and it is not a limitation of current models. It is arithmetic.

When a picture is reduced, information is destroyed. Four pixels become one, and the one that remains cannot distinguish between the many different sets of four that would have averaged to it. Enlarging is therefore not a calculation with a right answer; it is a choice among all the large pictures that would shrink to the small one you have.

A classical resampler picks the smoothest of those. A trained model picks the most likely one, judged by what pictures generally look like. That is a much better choice for making something look good. It is still a choice, and the model has no way of knowing which of the candidates was actually in front of the lens.

Researchers demonstrated this vividly with face upscaling: the same downsampled photograph, run through the same class of model with different settings, produces different people. Each result is internally consistent and sharp. Only one of them, at most, is the person.

What this means in practice

Good uses:

  • A photograph for a poster, a banner, a print or a product page, where it has to look right at a size it was not shot for.
  • An old family photo scanned at too low a resolution, to be printed and framed.
  • A supplier’s product image that is too small for a marketplace’s minimum.
  • A screenshot of a photograph for a presentation.
  • Texture and material in general: fabric, stone, foliage, food, fur.

Bad uses:

  • Anything evidential. A number plate, a face in a security still, a document, a serial number, a signature. The model will produce something legible and there is no reason to believe it.
  • Medical or scientific imagery, where an invented structure is worse than no structure.
  • Anything where a client will compare the enlargement with the real thing.

Uses where a classical filter is better:

  • Logos, icons, wordmarks, charts, diagrams, line drawings, screenshots of interfaces, anything with text in it. A model trained on photographs will add grain to a flat fill and a slight wobble to a clean curve, because that is what photographs do. Enlarging a logo without blur covers the right approach for these.

The honest failure modes

Knowing how a model fails is more useful than knowing that it can fail.

Waxiness. At high factors the model’s idea of skin overwhelms the subject’s actual skin, and faces look like moulded plastic. Lower the factor.

Invented text. Small lettering becomes crisp lettering that says something slightly different. This is the most dangerous failure because it is the most convincing.

Repeated texture. Areas of the picture with no real structure get given structure, and because the model has one idea of what that structure looks like, it repeats. Skies and out of focus backgrounds are where this shows.

Ringing at hard edges. Where the picture has a genuinely flat boundary, such as a logo on a photograph, the model may put a faint texture band along it.

Amplified noise. A noisy phone photo hands the model grain, and the model faithfully enlarges it into texture. On a very noisy original, a gentle denoise first is worth trying.

See the guess for yourself

  1. Take a sharp photograph you have at full size.
  2. Reduce it to a quarter of its width, so a 2000 pixel photo becomes 500 pixels. Save that.
  3. Enlarge the small version back to 2000 pixels with a model.
  4. Put the result beside the original at 100 percent.

You should see: A result that looks convincingly sharp, and that differs from the original in every fine detail: individual hairs in different places, fabric weave with a different pattern, small text that is legible and wrong.

If you do not: If they look identical, you did not zoom in far enough. Fit view hides all of this.

This is the single most useful exercise in this guide, because you have the ground truth. Everywhere else you do not, and it is easy to mistake “sharp” for “correct”.

Diffusion models change the trade, not the rule

The newest upscalers are not ESRGAN-style networks at all. They are diffusion models: the same machinery behind image generators, steered by your small picture instead of by a text prompt. They produce results that are startlingly detailed, often more convincing than anything a residual network manages.

They also make the underlying problem more obvious rather than less. A diffusion model is explicitly a sampler: it draws one plausible picture from a distribution of plausible pictures. Run it twice on the same input with different random seeds and you get two different answers, both sharp, both consistent with your file, and different from each other in every small detail. That is not a bug in the implementation, it is what the method is.

So the rule does not change with the technology. More capable models produce better looking guesses. They do not turn a guess into a measurement, and the gap between “looks like a photograph of this” and “is a photograph of this” is exactly as wide as it ever was.

Where the model runs, and why that matters

Most upscalers are services. You upload, a server does the work, you download. That is a reasonable engineering decision and it has a consequence: your picture is on somebody else’s computer, under terms you probably did not read. Several popular services reserve broad rights over what is sent to them. Is it safe to upload photos online goes through what the common tools actually say in their terms.

It does not have to work that way. ESRGAN-slim is small, about 900 KB of weights per scale, and a browser can run it on your graphics chip through WebGL. The picture stays on your device and the model comes to it. That is how the upscaler on this site works, and it is also why it offers a classical engine beside the model: once you are running both locally, giving you the choice costs nothing.

What to tell a client, in one sentence

If you hand an upscaled picture to somebody else, say so. The sentence that covers it without turning into a seminar is something like: “this image was enlarged from a smaller original, so fine detail has been reconstructed rather than photographed.” That is accurate, it is short, and it means nobody is surprised later when they compare the print with the product.

The same sentence belongs in an asset library. A file named product-hero-4x.jpg sitting next to product-hero.jpg tells the next person nothing about which one is a record. Putting the factor in the file name, as this site’s upscaler does automatically, is a small habit that saves a real argument.

A reasonable policy

Use the model when the picture is photographic and the output is for looking at. Use a classical filter when the picture is graphic or the output is for relying on. Never present an upscaled image as evidence of anything. And when it matters, say which engine you used and by what factor, because sooner or later somebody will put the enlargement next to the original and ask why the hairs moved.

The free way

Try both engines on the same picture

The upscaler here runs the model on your own graphics chip and offers the classical Lanczos engine beside it, so you can enlarge the same picture both ways and compare them at 100 percent with a draggable divider. Nothing is uploaded: the weights come down to your device once and are cached, and the picture never leaves the tab.

Your photo never leaves this device.Open the upscaler

Questions

Does AI upscaling actually add detail?

It adds detail that looks right, drawn from what the model has learned about pictures in general. It does not retrieve detail from your picture, because that detail is not in the file. The result genuinely looks more detailed and is genuinely not a record of what was in front of the lens at that scale.

Can AI enhance a face from CCTV?

No, not in the sense people mean. A model will produce a sharp, convincing face, and there is no reason to believe it is the right one. Researchers have shown the same low-resolution input producing different plausible faces depending on the model, which is exactly what you would expect from a system that guesses.

Is AI upscaling lossless?

No. Nothing about it is lossless. The output is a new picture that is consistent with the input, not a larger copy of it. If you need a faithful enlargement, use a classical resampler, which cannot invent anything.

Why do faces sometimes look plastic after upscaling?

Because the model is drawing its idea of skin rather than your subject's skin, and at high factors that idea dominates. Reducing the factor usually fixes it. So does using the model on the photograph and then blending it with a classical enlargement of the same picture.

Which model do the free upscalers use?

Most use some member of the ESRGAN family or a diffusion model on a server. This site uses ESRGAN-slim, which is MIT licensed, small enough to download in under a megabyte and good enough to run in a browser tab on your own graphics chip.

Is it safe to upload a photo to an upscaler?

It depends entirely on the service, and most of them do upload, because a big model needs a server. Read the terms: several reserve the right to keep or use what you send. If the picture is private, choose a tool that runs on your device.

Tags
UsesUpscalingResolutionPrivacy
The getPNG team

We build the free background remover this site runs on, which works inside your browser so your photos are never uploaded. These guides come from the questions people ask while cutting pictures out, and every method in them was checked against the app it describes.

Linking to this guide?

If it helped with a tutorial, a class page or an answer you are writing, a link is the kindest way to say so. Here it is, ready to paste.

<a href="https://getpng.app/guides/what-ai-upscaling-can-and-cannot-do">What AI upscaling can and cannot do</a>
What AI upscaling can and cannot do. getPNG, 23 Sept 2026. https://getpng.app/guides/what-ai-upscaling-can-and-cannot-do

Keep reading

All guides