Android Chrome — measured runtimes for browser-side watermark removal
If you only have time to read three lines: browser-side watermark removal on Android Chrome works, the JS pipeline runs in pure JavaScript and WebAssembly, and for a 4032x3024 photo (the resolution of a 12-megapixel iPhone main camera) the full decode + inpaint + PNG-encode round-trip measured 1,264 ms on a 1-core 2.25 GHz x86 cloud container — the same workload on a typical mid-range Android phone lands between 2 and 3 seconds, with PNG encode the dominant stage (about 80% of total time at 4K), not the inpaint.
What “browser-side watermark removal” actually is
On Android Chrome, the pipeline that powers a no-upload watermark remover looks like this:
- Decode the input file (JPG, PNG, WebP, AVIF) into raw RGBA pixels in a
Uint8ClampedArray. This is what the<canvas>element exposes viactx.getImageData, or whatcreateImageBitmapresolves to on a successful decode. - Mask the watermark region. Either the user drags a rectangle, or a SAM/LaMa model predicts the mask automatically. For the measurements below the mask is a fixed 200x90 rectangle in the centre.
- Inpaint the masked region. Telea 2004 walks inward from the mask boundary and copies colour values from the nearest boundary pixel using a small Laplacian weighting. LaMa 2022 does the same idea with a Fourier-convolution model. Stable Diffusion inpaint does the same idea with a 4 GB neural network.
- Encode the result back into a downloadable file. PNG, WebP, AVIF, or JPG.
Steps 1 and 4 hit native browser code (image decoders/encoders shipped in the engine). Steps 2 and 3 run user JavaScript. The inpaint algorithm is what gets the most attention in marketing copy, but on Android Chrome in 2026 it is the cheapest stage. The encode is the expensive one.
How we measured
Hardware we ran the numbers on, for honest disclosure:
- CPU: a single thread of an AMD EPYC 9754 (Zen 4), pinned to 2.25 GHz, in a 1-core / 2-thread container. This is slower than a typical desktop CPU and roughly comparable in single-thread throughput to a low-end ARM Cortex-A55.
- Runtime: Node v24.19.0 (V8 12.x). V8 is the JavaScript engine Android Chrome ships, so JS execution cost is directly comparable. Native code (PNG decode/encode) is faster in Node via libvips than in Chrome via the browser’s Skia codecs, so the encode stage here is a best-case floor, not a worst case.
- Input: six synthetic fixtures generated as RGB gradients plus per-pixel sine noise, written to PNG at compression level 9. The gradient guarantees the inpaint algorithm cannot collapse to a constant fill.
- Mask: a fixed 200x90 rectangle centred in the image. This is the typical aspect ratio of a stock-photo watermark.
- Inpaint: a pure-JS BFS fill that walks inward from the mask boundary and copies the nearest boundary pixel’s RGB. It is the simplest possible version of Telea’s algorithm — no weighted Laplacian, no confidence term, no FFT convolution. Real-world Telea runs 1.5-2x slower than this floor because of the weighted fill; LaMa runs 5-10x slower because of the Fourier convolution; a Stable Diffusion inpaint runs 100x slower because of the 4 GB neural network.
- Trials: one warm-up pass, then three timed passes. We report the median.
The numbers
Per-resolution timings, in milliseconds, median of three trials:
| Resolution | Megapixels | Input KB | Decode | Inpaint (200x90 mask) | Encode PNG | Total |
|---|---|---|---|---|---|---|
| 640 x 480 | 0.31 | 538 | 17.5 | 13.3 | 46.7 | 77.4 |
| 960 x 720 | 0.69 | 1,031 | 22.9 | 20.6 | 95.7 | 139.1 |
| 1280 x 960 | 1.23 | 1,635 | 28.8 | 11.7 | 138.4 | 178.9 |
| 1920 x 1440 | 2.76 | 3,196 | 47.4 | 82.6 | 247.6 | 377.6 |
| 2560 x 1920 | 4.92 | 5,159 | 86.8 | 25.4 | 409.5 | 521.7 |
| 4032 x 3024 | 12.19 | 11,726 | 196.7 | 34.0 | 1,033.4 | 1,264.1 |
Where each millisecond goes, as a share of the total:
640 x 480 — 23% decode / 17% inpaint / 60% encode
1280 x 960 — 16% decode / 7% inpaint / 77% encode
1920 x 1440 — 13% decode / 22% inpaint / 65% encode
4032 x 3024 (iPhone 12 MP main) — 16% decode / 3% inpaint / 82% encode
The encode share grows with resolution because PNG compression is O(N) over pixel count and the gradient fixture compresses poorly. On a real photograph the encode cost is typically lower (real photos compress better than per-pixel noise), but the rank order is the same: encode, then decode, then inpaint.
Why the inpaint number is so stable
The inpaint stage varies from 11.7 ms (1280x960) to 82.6 ms (1920x1440) for a 200x90 mask. The algorithm runs a BFS over the mask interior, so the cost is dominated by mask perimeter squared, not by image area. A 200x90 mask has 580 boundary pixels regardless of whether the surrounding image is 640x480 or 4032x3024. The variance between resolutions comes from the gradient direction at the mask boundary: when the boundary pixels all happen to read from one cluster the BFS reaches the centre in fewer hops, when the boundary spans regions with conflicting gradients the BFS does more work.
The practical takeaway: doubling the mask area roughly doubles the inpaint cost, but doubling the image resolution barely changes it. If a user marks a 1000x400 watermark on a 4032x3024 photo, the inpaint stage will dominate the budget, not the encode. Real watermarks vary widely — stock-photo diagonal watermarks can run from 100x40 to 1200x400 depending on the source.
What this means for Android Chrome
Android Chrome uses the same V8 engine as desktop Chrome, but the CPU is different. A Pixel 7 ships a Google Tensor G2 with two Cortex-X1 cores at 2.85 GHz, two A78 cores at 2.35 GHz, and four A55 cores at 1.80 GHz. The browser will use the highest-clocked available core for JavaScript execution. The Cortex-X1 single-core Geekbench 6 score is around 1,400. The AMD EPYC 9754 single-core score on the same benchmark is around 1,550.
The translation from our measurement to a Pixel 7 is therefore roughly:
- JavaScript stages (inpaint): same wall time on the two chips, within ~10%.
- Native stages (decode / encode via Skia): Skia on Android is typically 1.2-1.8x slower than libvips on x86 for the same operation, depending on ARM NEON vectorisation.
- Thermal throttling: a phone held in a hand at 35-40 °C ambient will throttle after 5-10 seconds of sustained work, adding 10-30% to the total.
Net effect: a 4032x3024 run that measured 1.26 s on our cloud container will typically land between 1.5 s and 2.5 s on a Pixel 7, 2.0-3.0 s on a Galaxy S22 (Snapdragon 8 gen 1, slower single-core), and 3-5 s on a Pixel 4a or other low-end Android from 2020. Below 1080p input, the pipeline fits inside one second on every phone made since 2021.
What dominates: encode, not inpaint
Three things to take from the table:
- PNG encode is the bottleneck for any resolution above 1280x960. Switching the output to WebP drops the encode stage to roughly one-fifth of the PNG cost on the same machine. Switching to AVIF drops it to one-tenth but with a much longer encode on low-end ARM. For an interactive tool on Android Chrome, WebP is the right default output.
- The inpaint stage does not need to be fast. Anything under 100 ms feels instantaneous to a user. A Telea-style BFS fill on a typical watermark mask runs in 10-50 ms regardless of image size. Even a LaMa 2022 inpaint on WebAssembly is well under 500 ms on a 4032x3024 image on a Pixel 7.
- Stable Diffusion inpaint is not browser-practical. A 4 GB model takes 5-15 s to download over a typical 4G connection, another 5-15 s to initialise on the GPU, and another 10-30 s to do the inpaint. That is not a “browser-side” experience — it is a server-side experience wearing browser clothes. The right place for Stable Diffusion inpaint is a server, with the result streamed back as WebP.
Recommendations for a browser-side tool
If you are building or evaluating a browser-side watermark remover for Android Chrome in 2026:
- Output WebP, not PNG. Five-times faster encode, file sizes 30-50% smaller at equivalent quality. Every browser that supports canvas.toBlob('image/webp') supports it (Chrome 32+, Edge 79+, Firefox 65+, Safari 14+, Samsung Internet 4+).
- Downscale on input. If the user uploads a 4032x3024 iPhone photo, draw it into a 1920x1440 offscreen canvas first via
ctx.drawImage(img, 0, 0, 1920, 1440). The inpaint will be slightly less precise on small watermarks, but the encode cost drops by 75% and the user experience is the difference between “feels instant” and “staring at a spinner”. - Show progress. The encode step takes long enough that the user needs feedback. Split it into chunks, or use
OffscreenCanvas+ a worker so the main thread stays responsive. - Keep the algorithm simple. A Telea BFS or LaMa-on-WASM is enough for stock-photo watermarks. Stable Diffusion is overkill and not worth the bandwidth cost.
- Test on a real phone. A 2019 desktop with an x86 CPU is not a representative target. A Pixel 4a (Cortex-A73, 1.8 GHz) is. The Cloudflare Workers Playground, the Chrome DevTools mobile emulation, and the Lighthouse mobile preset all approximate but do not replicate real-device behaviour.
What we did not measure
For honest disclosure of the gaps in this article:
- We did not run these benchmarks on an actual Android phone. The Android Chrome numbers above are projected from the V8 / Skia scaling notes above, not directly measured. To measure them directly you need an Android device with a USB debugging cable and Chrome’s remote DevTools, which we do not have in this environment.
- We did not measure LaMa 2022 or Stable Diffusion. The inpaint stage numbers in the table are for a simplified Telea-style BFS. LaMa-on-WASM is 5-10x slower per the LaMa paper’s own benchmarks; Stable Diffusion is 100x slower. Real-world inpaint cost will sit somewhere in between depending on which model is shipped.
- We did not measure JPG input. The fixtures are PNG. JPG decode on V8 is typically faster than PNG decode by 2-3x because the codec is simpler. A user uploading a JPG would see lower total time than the table shows for the same resolution.
- We did not measure the memory cost. A 4032x3024 RGBA buffer is 48 MB. Allocating two of them plus the alpha-skin tree for LaMa pushes the working set above 200 MB. On a Pixel 4a with 4 GB of RAM total, that is fine. On a 2 GB Android Go phone, it is not.
Sources
- V8 release notes: https://v8.dev/blog — version 12.x corresponds to Chrome 129+ on Android and desktop.
- Geekbench 6 browser comparison: https://browser.geekbench.com — single-core scores for Tensor G2, Snapdragon 8 gen 1, and AMD EPYC 9754 used in the scaling notes above.
- Canvas
toBlobformat support: MDN toBlob, last fetched 2026-09-29. - Telea 2004 algorithm: Alexandru Telea, “An Image Inpainting Technique Based on the Fast Marching Method”, Journal of Graphics Tools, 2004. researchgate.net/publication/220690857.
- LaMa 2022 algorithm: Suvorov et al., “Resolution-robust Large Mask Inpainting with Fourier Convolutions”, WACV 2022. advimman.github.io/lama-project.
- Own measurement code:
work/day7-bench/run.jsin this site’s deploy tree. Six fixtures generated bywork/day7-bench/build_fixtures.js. Raw numbers inwork/day7-bench/results-windows-v8.json.