Skip to main content

OpenAI’s Sites feature will take a sentence and give you back a hosted website on its own address. For a founder who needs a landing page by Friday, that is an attractive trade. The question we had was narrower and more boring: what does the server actually send back once the site is live?

So we measured it. We assembled 89 publicly reachable ChatGPT-hosted sites, probed 86 of them that were still serving on 3 October 2026, and looked at the response layer rather than the design. The HTML those sites produce is largely fine. The HTTP around it is not.

TL;DR

  • 46 of 86 live ChatGPT-hosted sites return HTTP 200 for a URL that does not exist. We asked for /zzqq-nonsense-7351, /wp-admin/ and /.env on every site and got 200 and an HTML page back from the same 46. The other 39 correctly returned 404.
  • 44 of 86 serve /robots.txt with status 200 and Content-Type: text/html, meaning the file is the homepage, not crawl directives. Only 9 of 86 served a parseable robots.txt.
  • 80 of 86 send none of the six standard security response headers. 82 of 86 can be framed by any site on the internet, and 1 of 86 sends HSTS.
  • Every live site we found uses the form site.workspace.chatgpt.site. The middle label is your workspace, it is permanent, and it is the same on every site you publish. In our wider sample of 3,397 hostnames, seven of those workspace labels reconstruct a Gmail address verbatim.
  • Someone is already mapping the namespace. We counted 3,400 public scans of *.chatgpt.site hostnames submitted to urlscan.io’s API in the eleven hours to 07:38 UTC on 3 October 2026, a sustained rate of roughly 250 an hour.

How we built the sample

There is no directory of published ChatGPT Sites, which is itself part of the story. We used urlscan.io’s public search index to collect 3,397 distinct chatgpt.site hostnames, filtered to those that returned 200 at scan time, and ended up with 89 live public sites across 85 distinct workspaces. Three had moved behind authentication by the time we probed, leaving 86.

We checked the homepage, then six fixed paths on each host, then the workspace root, each with an ordinary desktop browser user agent. Control probes against known-good OpenAI-operated sites returned 200 throughout, so the failures below are not us being rate limited.

Finding one: half the fleet has no 404

This is the result that matters most and it is the easiest to verify yourself. On 46 of the 86 sites, every path we invented returned HTTP/2 200 with content-type: text/html. Not a redirect, not a soft error page with a 404 status. A 200.

That is the classic single-page-application catch-all, and on a dashboard behind a login it is harmless. On a public marketing site it is not. Search engines treat a 200 response as a real, indexable page. A site with no 404 has an unbounded URL space in which every address returns content, which is the textbook definition of a soft 404 and of mass duplicate content. Only 9 of the 86 sites carry a canonical tag, so there is nothing telling a crawler which of those infinite addresses is the real one.

The practical consequence is mundane and expensive. Anyone can mint links to your domain that resolve, including links a competitor or a spammer generates. Your crawl budget gets spent on them. Your analytics fill with paths you never built.

Note the split: 39 sites do return a proper 404. Two serving behaviours exist on the same platform and nothing in the product surface tells you which one you have been given.

Finding two: your robots.txt is an HTML page

We fetched /robots.txt from all 86. Thirty-three returned 404, which is fine and conventionally means “crawl everything”. Fifty-three returned 200, and 44 of those 53 returned HTML. We confirmed by hand on multiple hosts: status 200, content-type: text/html, body beginning <!doctype html>.

A crawler asking for robots.txt and receiving a 200 with an HTML body does not get your rules. It gets a document it cannot parse, and the standard behaviour is to proceed as if no restrictions exist. The same pattern applies to /sitemap.xml: 52 sites answered 200 and 46 of those answers were HTML.

If your business depends on controlling what gets crawled, on staging pages that must not be indexed, or on submitting a sitemap, more than half this platform cannot do it. That is not a configuration you have got wrong. There is no configuration.

Finding three: the response headers are empty

We checked six headers: Content-Security-Policy, Strict-Transport-Security, X-Frame-Options, X-Content-Type-Options, Referrer-Policy and Permissions-Policy.

80 of 86 sites sent none of them. One site sent HSTS. Five sent a CSP, and in each case it was clearly authored by the generated application rather than supplied by the platform. Taking X-Frame-Options and frame-ancestors together, 82 of 86 sites can be loaded in an iframe by any origin, which is the precondition for clickjacking.

Twelve of the 86 sites render an HTML <form>, so they are collecting something from visitors. Ten of those twelve have no CSP at all. Add the platform’s own “Sign in with ChatGPT” and built-in storage features, and you have sites taking authenticated user input with no declared script policy and no framing protection.

One mitigating detail: plain HTTP does redirect to HTTPS with a 302. Without HSTS, though, that first unencrypted request still happens.

Finding four: the hostname is your identity

Every live site used site.workspace.chatgpt.site. The workspace label is stable across everything you publish, so one shared link permanently associates all your other sites with each other. All 86 workspace roots returned 404, so the root itself does not enumerate your sites, but the label is still a durable public identifier attached to a person or a company.

Users who choose their own workspace label get no warning that it becomes a public hostname. In our 3,397-hostname sample, 462 labels were self-chosen rather than auto-generated, and seven of those reconstruct a working Gmail address character for character. Others are plainly company names.

Access control leaks too. An OpenAI-operated site that is gated returns 401, and that 401 page carries the Site’s real name in its <title> element. An invented site name under the same workspace returns 404. A 401 therefore confirms a restricted site exists and tells you what it is called, without any credential.

Finding five: the namespace is being swept now

The auto-generated default label follows an adjective-noun-four-digit pattern. Across our sample we counted 378 distinct adjectives and 225 distinct nouns, giving a keyspace of roughly 850 million. Nobody is guessing their way through that at a few hundred requests an hour, which means the hostnames being probed were obtained, not derived.

And they are being probed. Every one of the 3,400 scans we counted in that eleven-hour window was submitted through urlscan.io’s API with public visibility, at a steady rate with one burst of 899 in a single hour. 3,382 returned 404. We cannot attribute the activity, and the choice of public visibility means the results form a free, queryable map of the namespace for anyone who wants it. A platform this young being swept this systematically is worth knowing about before you put a client’s brand on it.

What we would actually do

Use it for what it is good at. Internal tools, a throwaway prototype, a one-weekend microsite, a client demo. The generated HTML is competent: 86 of 86 had a title, 77 had an h1, 74 had a meta description. The markup is not the problem.

Do not use it as the public front door of a business. If a page needs to rank, control crawling, collect customer data, or carry your brand under its own domain, it needs a host where you own the response headers and the 404. Point a custom domain at it and you inherit every finding above on your own domain name rather than OpenAI’s, which makes the SEO exposure worse, not better.

And if you already have one of these live, run three commands today: request a path that does not exist and check the status code, fetch /robots.txt and check the content type, and run curl -I on the homepage and read what comes back. Five minutes tells you which half of the fleet you are in.

REPTILEHAUS builds and operates production web platforms, and we spend a lot of our time on exactly this layer: headers, caching, crawl control and the unglamorous HTTP semantics that decide whether a site performs. If you have prototyped something on an AI site builder and now need it to hold up as a real product, get in touch.

Measurements taken 3 October 2026. Sample: 86 publicly reachable ChatGPT-hosted sites across 85 workspaces, plus 3,397 hostnames observed in public scan data.

📷 Photo by Vadim Babenko (@vakerbv) on Unsplash