# How PhotoAIBench tests AI photo tools

> The rules behind every PhotoAIBench test. How the public-domain and CC0 test photos are chosen and licensed, how an upscaler is measured, what PSNR, SSIM and LPIPS mean in plain words, how commercial tools are tested within their terms, and the background-removal and restoration tests that are planned but not yet run.

By PhotoAIBench. Updated 7 October 2026. Canonical URL: https://photoaibench.com/methodology/

## Rules

1. **Measured, or not published.** No score appears on this site until it has been measured with the method on this page, and every published result comes with its data and the code that produced it.
2. **Test photos must be free to use.** Public domain or CC0 only, each with its source and license recorded. No research-only datasets, nothing with a non-commercial or share-alike condition, nothing whose status is unclear.
3. **No identifiable living people.** Tests upload photos to other companies' services and publish crops of what comes back. We won't do that with someone's face without their consent, so the portraits are historical.
4. **Terms first.** Before a commercial tool is used, its terms of service are read and the relevant clauses recorded. A tool whose terms forbid benchmarking, publishing results, competitive analysis or automated access is not tested.
5. **Free tiers, no accounts, no workarounds.** Commercial tools are tested only through what anyone can use without an account or a payment card, through their normal web page, one image at a time with at least a minute between images. We never get around a captcha, challenge, quota, rate limit or paywall; where a tool stops us, we stop and say so. Testing with an account or a paid plan is the site owner's decision, in his own name, and would be stated.
6. **Vendors see their results first.** A result that names a commercial tool is sent to that vendor, and published no sooner than seven days later, with any reply.
7. **Same code for everyone.** Every output, including from free open-source methods, is scored by the same script against the same reference images.

## The test photos

16 photos in four groups: 4 portraits, 4 landscapes, 4 with fine texture or lettering, and 4 old or damaged photos for restoration. The first three groups (12 photos) are scored in the upscaler test; the damaged ones are for side-by-side restoration examples only, because a restoration has no single correct answer to score against.

<section class="pricing-data" data-block="data" markdown="1">

<div class="table-wrap" markdown="1">

| Photo | Group | Source and license | Original size |
|---|---|---|---|
| [Guide at Little Norway, Blue Mounds, Wis.](https://www.loc.gov/item/2017878934/) | Portrait | No known restrictions on publication. Arthur Rothstein (1915-1985), photographer, for the US Farm Security Administration; Library of Congress, Prints & Photographs Division, Farm Security Administration/Office of War Information Color Photographs | 5063 × 6657 |
| [Garage mechanic near Newark, N.J. Badge denotes member of Office...](https://www.loc.gov/item/2017878857/) | Portrait | No known restrictions on publication. Marjory Collins (1912-1985), photographer, for the US Office of War Information; Library of Congress, Prints & Photographs Division, Farm Security Administration/Office of War Information Color Photographs | 6162 × 7981 |
| [Walter Francis White by Clara Sipprell](https://commons.wikimedia.org/wiki/File:Walter_Francis_White_by_Clara_Sipprell.jpg) | Portrait | CC0 (Creative Commons Zero, Public Domain Dedication) Clara Sipprell (photographer, circa 1950); digital restoration by Adam Cuerden; National Portrait Gallery, Smithsonian Institution (NPG.82.197) | 3520 × 4536 |
| [Alexander Graham Bell 1895 NPG 77 363](https://commons.wikimedia.org/wiki/File:Alexander_Graham_Bell_1895_NPG_77_363.jpg) | Portrait | CC0 (Creative Commons Zero, Public Domain Dedication) Unknown photographer; National Portrait Gallery, Smithsonian Institution (NPG.77.363) | 4726 × 7001 |
| [Parc national de la Jacques-Cartier, Quebec, Canada 22](https://commons.wikimedia.org/wiki/File:Parc_national_de_la_Jacques-Cartier,_Quebec,_Canada_22.jpg) | Landscape | CC0 (Creative Commons Zero, Public Domain Dedication) Wilfredor; Wilfredor, Wikimedia Commons, CC0 | 8256 × 5504 |
| [Mount Ida chain Messara plain from Phaistos Crete Greece](https://commons.wikimedia.org/wiki/File:Mount_Ida_chain_Messara_plain_from_Phaistos_Crete_Greece.jpg) | Landscape | CC0 (Creative Commons Zero, Public Domain Dedication) Jebulon; Jebulon, Wikimedia Commons, CC0 | 4479 × 2730 |
| [Rio de Janeiro skyline and Sugarloaf Mountain at sunset, Brazil 3](https://commons.wikimedia.org/wiki/File:Rio_de_Janeiro_skyline_and_Sugarloaf_Mountain_at_sunset,_Brazil_3.jpg) | Landscape | CC0 (Creative Commons Zero, Public Domain Dedication) Wilfredor; Wilfredor, Wikimedia Commons, CC0 | 8043 × 4919 |
| [Sunset Key and Sailboats, Florida, 2025](https://commons.wikimedia.org/wiki/File:Sunset_Key_and_Sailboats,_Florida,_2025.jpg) | Landscape | CC0 (Creative Commons Zero, Public Domain Dedication) Julian Lupyan; Julian Lupyan, Wikimedia Commons, CC0 | 6964 × 3626 |
| [Hotel Post, Inn sign, Murnau, Bavaria, Germany](https://commons.wikimedia.org/wiki/File:Hotel_Post,_Inn_sign,_Murnau,_Bavaria,_Germany.jpg) | Texture or text | CC0 (Creative Commons Zero, Public Domain Dedication) Jebulon; Jebulon, Wikimedia Commons, CC0 | 2988 × 4248 |
| [Posted Nature Sanctuary sign, Capen Hill Nature Sanctuary](https://commons.wikimedia.org/wiki/File:Posted_Nature_Sanctuary_sign,_Capen_Hill_Nature_Sanctuary.jpg) | Texture or text | CC0 (Creative Commons Zero, Public Domain Dedication) Peter Cooper Jr.; Peter Cooper Jr., Wikimedia Commons, CC0 | 6000 × 4000 |
| [Portrait of a Carpathian Lynx](https://commons.wikimedia.org/wiki/File:Portrait_of_a_Carpathian_Lynx.jpg) | Texture or text | CC0 (Creative Commons Zero, Public Domain Dedication) Wilfredor; Wilfredor, Wikimedia Commons, CC0 | 7094 × 4998 |
| [Golden Chapel, Recife, Pernambuco, Brazil](https://commons.wikimedia.org/wiki/File:Golden_Chapel,_Recife,_Pernambuco,_Brazil.jpg) | Texture or text | CC0 (Creative Commons Zero, Public Domain Dedication) Wilfredor; Wilfredor, Wikimedia Commons, CC0 | 6473 × 6255 |
| [[Unidentified copy of a daguerreotype portrait of a woman]](https://www.loc.gov/item/2004664223/) | Old or damaged | No known restrictions on publication. Mathew B. Brady studio (approximately 1823-1896); Library of Congress, Prints & Photographs Division, Daguerreotypes Collection | 3295 × 4096 |
| [[Unidentified man, head-and-shoulders portrait, facing front]](https://www.loc.gov/item/2004664536/) | Old or damaged | No known restrictions on publication. F. (Francis) Grice, photographer; Library of Congress, Prints & Photographs Division, Daguerreotypes Collection | 3746 × 4697 |
| [[Half-length portrait of an African American woman wearing a hat...](https://www.loc.gov/item/2010650827/) | Old or damaged | No known restrictions on publication. Unidentified photographer; Library of Congress, Prints & Photographs Division, Gladstone Collection of African American Photographs | 3051 × 4383 |
| [Colorado River. Grand Canyon, Tapeets Creek. (Note: it appears...](https://commons.wikimedia.org/wiki/File:Colorado_River._Grand_Canyon,_Tapeets_Creek._(Note,_it_appears_that_the_glass_negative_may_have_broken_or_cracked_all..._-_NARA_-_518020.jpg) | Old or damaged | Public domain (work of the US federal government) Elias Olcott Beaman, James Fennemore or John Karl Hillers (Powell survey photographers); U.S. National Archives and Records Administration (NAID 518020) | 3000 × 1871 |


</div>

The full manifest, with each file's download address, SHA-256 hash and the reasons each photo is free to use: [photo-samples.yaml](/tools/image-quality-compare/photo-samples.yaml).

</section>

**How they were chosen.** We searched Wikimedia Commons for photos dedicated to the public domain under CC0 (mostly among its featured and quality pictures, which are checked by volunteers for technical quality), and the Library of Congress and the US National Archives for public-domain photos. We picked photos that are sharp at the pixel level and that cover what upscalers struggle with: skin and hair, foliage, water, thin lines such as masts, dense ornament, fur, and lettering. Then we checked each file's license on its own page and recorded it word for word.

**What the set does not cover yet.** No modern digital portrait (see rule 3), no low-light or noisy phone photos, no screenshots or illustrations, and only twelve scored photos. Twelve is enough to see large differences between methods, not small ones.

**Reproducing the set.** The photos are not stored in our code repository. A script downloads each from its source and checks its hash, so you get exactly the same bytes or an error.

## How an upscaler is tested

The question an upscaler answers is "what would this small image look like if it had been taken at a higher resolution?" To score that, you need the higher-resolution answer. So the test starts from a large photo, makes a small copy, asks each method to enlarge the small copy, and compares the result with the large original.

1. **Reference.** Each photo is converted to sRGB (two are stored in ProPhoto RGB and one in Adobe RGB), film borders are cropped from the two Kodachrome transparencies, and the photo is resized so its long side is 2,000 px with a Lanczos filter. Shrinking a 20 to 50 megapixel original to about 2.7 megapixels averages away sensor noise, film grain and JPEG artifacts, so the reference is clean and sharp at the pixel level.
2. **Input.** The reference reduced by exactly 4x with Pillow's bicubic filter (500 px on the long side). This "bicubic degradation" is the standard in super-resolution research. It is an easy case: real small images also carry blur, noise and compression, which we plan to add as a second, harder test.
3. **Upscale.** Each method enlarges the input 4x, back to the reference's exact size. Local methods run on our computer. Commercial web tools get the input file through their own web page.
4. **Score.** Each output is compared with the reference with three metrics, below. No border is cropped: what a tool does at the edges is part of its output.
5. **Look.** Crops of the same region from every output are published side by side, chosen by a fixed rule (the 128 × 128 region with the most detail in the reference), not by eye.

## The metrics, in plain words

**PSNR (peak signal-to-noise ratio)** asks: on average, how far is each pixel value from the reference? It is the mean squared error put on a logarithmic scale, in decibels. Higher is better. A gain of 6 dB means the typical error is half as large.

```
PSNR = 10 × log10(255² / mean((reference − output)²))
```

**SSIM (structural similarity)** asks: in each small neighbourhood, do the two images have the same brightness, the same contrast, and the same pattern? It compares 11 × 11 windows and averages the result. 1 means identical. It follows edges and textures better than PSNR, because a pattern that is correct but slightly darker scores well.

```
SSIM = ((2·μx·μy + C1) · (2·σxy + C2)) / ((μx² + μy² + C1) · (σx² + σy² + C2))
```

**LPIPS (learned perceptual image patch similarity)** asks: do the two images look alike to a neural network trained to see the way people judge similarity? Both images go through a pretrained image network (AlexNet), and LPIPS measures the distance between the features it extracts, with weights fitted to thousands of human judgements. Lower is better; 0 is identical.

PSNR and SSIM on our pages use exactly the settings of our free [PSNR and SSIM calculator](/psnr-ssim-calculator/) (luma Y = 0.299 R + 0.587 G + 0.114 B; SSIM per Wang et al. 2004 with an 11 × 11 Gaussian window, σ 1.5), so you can check any number yourself. LPIPS uses the reference `lpips` package (version 0.1.4, AlexNet, v0.1).

### What these numbers can and can't tell you

- **They measure faithfulness, not beauty.** An upscaler can only guess the missing detail. PSNR and SSIM reward the safest guess, which is the average of all plausible answers and therefore a little blurry. Generative (GAN) upscalers draw sharp, plausible detail that isn't exactly where the real detail was, and PSNR punishes that even when people prefer it. LPIPS sides with people more often, but it is a network's opinion, not a person's.
- **They can't see invented content as a mistake in kind.** Wrong letters on a sign, an extra eyelash or a changed facial expression may cost only a fraction of a decibel. That is why we publish crops, and why you should look at text and faces yourself.
- **Small differences mean nothing.** Twelve photos and one degradation type can separate methods that differ by several decibels. A 0.2 dB gap between two tools is noise for a test this size.
- **The test is about enlarging clean, small photos.** It says nothing about denoising, deblurring, JPEG artifact removal or restoring damage, which need their own tests.

## Commercial tools

We read the terms and free tiers of 26 commercial and online upscalers before using any of them (rule 4). Most could not be tested: their terms forbid automated access or benchmarking, or their free tier needs an account. A pilot round with the tools that remained has been measured; it goes to the vendors first (rule 6) and will be published after their seven days, with the full list of tools considered and why each was or wasn't tested.

## Planned tests (not run yet)

**Status: planned. Nothing below has been measured.** The method is written down first so it can't be bent to fit the results.

- **Background removal.** For each test photo, a hand-made alpha mask (the "answer"), made by a person in an image editor at full resolution and published with the photos. Each tool's cut-out is scored by mask overlap (intersection over union), by accuracy along the boundary (a boundary F-score within a few pixels of the true edge), and by alpha error on hair and fur. Photos will include hair, fur, glass, a product shot and a busy background. Hand-making masks takes hours per photo, so this test starts with six photos.
- **Photo restoration.** The four damaged photos above, run through restoration tools, published side by side at full resolution with the original. No score: restoration invents what was lost, and there is no reference to compare with. We will describe what each tool changed (crack filled, scratch removed, face redrawn) and flag invented detail.
- **Harder upscaling.** The same photos, but with the small copy blurred, noised and JPEG-compressed before upscaling, which is closer to real old phone photos and web images.

## Data and code

- [photo-samples.yaml](/tools/image-quality-compare/photo-samples.yaml): every photo, its license and its hash.
- [upscaler-baselines.json](/tools/image-quality-compare/upscaler-baselines.json): every score of the local baseline methods, per photo.
- The scripts: [prepare](/tools/image-quality-compare/bench-prepare.py) (reference and input images), [upscale](/tools/image-quality-compare/bench-upscale-local.py) (local methods) and [score](/tools/image-quality-compare/bench-score.py) (PSNR, SSIM, LPIPS).

## Conflicts of interest

PhotoAIBench may earn affiliate commissions from some photo tools (see the [affiliate disclosure](/affiliate-disclosure/); none are active today). A commission never decides which tools are tested, how, or in what order results appear: tables are ordered by the measured score, and tools that pay nothing are tested the same way.

## Corrections

If a license, a source or a number on this site is wrong, the contact details are on the [about page](/about/). Corrections are dated on the page they affect.
