Extensions / Chrome Extension

Any web page, as a starter project.

SiteExtract saves the page you are viewing as clean HTML, CSS, images, fonts and design tokens, in one ZIP you can open and edit. Everything runs in your browser.

Live on the Chrome Web Store · Free

One ZIP, ready to open

Every reference is rewritten to a local path, so index.html opens straight from your disk. No build step.

siteextract-example.com-20260927/
├── index.html
├── css/
│   ├── styles.css
│   ├── inline.css
│   └── tokens.css
├── images/
├── fonts/
├── report.json
└── README.md

index.html

The page, with its links pointing at local files.

css/

The page's stylesheets in cascade order, @import rules inlined, and inline styles kept in their own file.

images/ and fonts/

Downloaded and deduplicated, so the same file served from two addresses is stored once.

tokens.css

The colours, type scale, spacing, radii and shadows the page actually uses, ranked by how often they appear.

report.json

Anything that could not be downloaded, and why.


The page you see, or the page as written

Snapshot Default

Saves the page as rendered: form state, lazy-loaded images, shadow DOM and styles added by scripts. Every script and inline handler is removed, leaving a static page you can edit.

Source

Downloads the original HTML again and keeps its scripts, saved to js/. Closer to how the site was built, though apps that load data at runtime may not work offline.


Four steps, no setup

1

Open a page

A layout you want to learn from, or a site you are rebuilding.

2

Open SiteExtract

Click the toolbar button, or press Ctrl Shift E.

3

Press Extract

Pick a mode, choose images, fonts and tokens, and watch the files come in.

4

Open the ZIP

Unzip it and open index.html. It works offline.


It tells you what it missed

Nothing silently dropped

A file that fails to download keeps its original URL, so the page still works online, and report.json lists it with the reason.

Tokens are a starting point

They are drawn from what the page renders, not a claim about the site's real design system. Near-identical values are merged and ranked by use.

Your project, your marks

A labelled HTML comment credits SiteExtract at the top of index.html. A visible badge is off by default. Both switch off in settings, and no advert ever goes inside the page.

Their work stays theirs

The project is for learning, prototyping and sites you have the right to copy. Text, images and fonts belong to their owners, and font licences rarely allow reuse.


There is nowhere to upload anything to

SiteExtract has no servers, no account and no telemetry. The project is built in your browser and saved only to your downloads.

Not on the page until you ask

There is no content script in the manifest. Nothing of SiteExtract exists on a page until you press Extract on it.

522 sites it will not read

Large platforms, and pages that routinely hold personal information such as webmail, cloud storage and banking. There is no override.

Passwords never saved

Password, card number and one-time code fields are emptied before the page leaves the tab.

Cookies stay on the page's site

Files from other domains are requested without cookies. Blocked sites and private network addresses are never requested at all.

Asked for, not assumed

Access to other domains is requested on your first extract, not at install. Decline it and those files keep their remote URLs.

No history

Only your settings are stored, on your device. There is no list of the sites you have extracted.

Read the full privacy policy

Common questions

What does it cost?

Nothing, and it stays that way. SiteExtract carries one house advertisement for the studio's own work at the foot of the popup, labelled "Ad", served from a list that ships inside the extension. One click turns it off for good.

Which browsers does it work in?

Chrome 116 and later, and the Chromium browsers that install Chrome extensions, such as Edge, Brave and Opera.

Why will it not extract the page I am on?

It is probably on the block list: a large platform, or a kind of page that routinely holds personal information. The popup names the site and says why. Sites people deploy their own work to, such as GitHub Pages, Vercel and Netlify, are never blocked.

Why did some files not download?

Usually because access to other domains was declined, so files on CDNs stayed remote. Grant it from the settings page and extract again. Every failure is listed in report.json with its reason.

Does it work on a page behind a login?

Yes. A snapshot saves the page as you see it, signed in, and it never leaves your device. Password fields are always emptied, but text in other fields is kept, so clear anything you do not want saved first.

Not answered here? Extension support has the troubleshooting, and what to include in a bug report.

Ready when you are

It installs in a click and costs nothing. While you are here: we also build websites, and we make more extensions like this one.