Any web page, as a starter project.
SiteExtract saves the page you are viewing as clean HTML, CSS, images, fonts and design tokens, in one ZIP you can open and edit. Everything runs in your browser.
One ZIP, ready to open
Every reference is rewritten to a local path, so index.html
opens straight from your disk. No build step.
siteextract-example.com-20260927/ ├── index.html ├── css/ │ ├── styles.css │ ├── inline.css │ └── tokens.css ├── images/ ├── fonts/ ├── report.json └── README.md
index.html
The page, with its links pointing at local files.
css/
The page's stylesheets in cascade order, @import rules inlined, and inline
styles kept in their own file.
images/ and fonts/
Downloaded and deduplicated, so the same file served from two addresses is stored once.
tokens.css
The colours, type scale, spacing, radii and shadows the page actually uses, ranked by how often they appear.
report.json
Anything that could not be downloaded, and why.
The page you see, or the page as written
Snapshot Default
Saves the page as rendered: form state, lazy-loaded images, shadow DOM and styles added by scripts. Every script and inline handler is removed, leaving a static page you can edit.
Source
Downloads the original HTML again and keeps its scripts, saved to js/. Closer to
how the site was built, though apps that load data at runtime may not work offline.
Four steps, no setup
Open a page
A layout you want to learn from, or a site you are rebuilding.
Open SiteExtract
Click the toolbar button, or press Ctrl Shift E.
Press Extract
Pick a mode, choose images, fonts and tokens, and watch the files come in.
Open the ZIP
Unzip it and open index.html. It works offline.
It tells you what it missed
Nothing silently dropped
A file that fails to download keeps its original URL, so the page still works online, and
report.json lists it with the reason.
Tokens are a starting point
They are drawn from what the page renders, not a claim about the site's real design system. Near-identical values are merged and ranked by use.
Your project, your marks
A labelled HTML comment credits SiteExtract at the top of index.html. A visible
badge is off by default. Both switch off in settings, and no advert ever goes inside the
page.
Their work stays theirs
The project is for learning, prototyping and sites you have the right to copy. Text, images and fonts belong to their owners, and font licences rarely allow reuse.
There is nowhere to upload anything to
SiteExtract has no servers, no account and no telemetry. The project is built in your browser and saved only to your downloads.
Not on the page until you ask
There is no content script in the manifest. Nothing of SiteExtract exists on a page until you press Extract on it.
522 sites it will not read
Large platforms, and pages that routinely hold personal information such as webmail, cloud storage and banking. There is no override.
Passwords never saved
Password, card number and one-time code fields are emptied before the page leaves the tab.
Cookies stay on the page's site
Files from other domains are requested without cookies. Blocked sites and private network addresses are never requested at all.
Asked for, not assumed
Access to other domains is requested on your first extract, not at install. Decline it and those files keep their remote URLs.
No history
Only your settings are stored, on your device. There is no list of the sites you have extracted.
Common questions
What does it cost?
Nothing, and it stays that way. SiteExtract carries one house advertisement for the studio's own work at the foot of the popup, labelled "Ad", served from a list that ships inside the extension. One click turns it off for good.
Which browsers does it work in?
Chrome 116 and later, and the Chromium browsers that install Chrome extensions, such as Edge, Brave and Opera.
Why will it not extract the page I am on?
It is probably on the block list: a large platform, or a kind of page that routinely holds personal information. The popup names the site and says why. Sites people deploy their own work to, such as GitHub Pages, Vercel and Netlify, are never blocked.
Why did some files not download?
Usually because access to other domains was declined, so files on CDNs stayed remote. Grant
it from the settings page and extract again. Every failure is listed in
report.json with its reason.
Does it work on a page behind a login?
Yes. A snapshot saves the page as you see it, signed in, and it never leaves your device. Password fields are always emptied, but text in other fields is kept, so clear anything you do not want saved first.
Not answered here? Extension support has the troubleshooting, and what to include in a bug report.
Ready when you are
It installs in a click and costs nothing. While you are here: we also build websites, and we make more extensions like this one.