To summarize a website, choose the exact webpage you need rather than an entire domain, record its title, publisher, author, URL, date, and access date, then provide the accessible page URL or selected text to a summarizer. Compare the result with the original headings and verify every important name, number, quotation, date, and conclusion. If the page is dynamic, private, paywalled, or extracted incompletely, work from text you are permitted to use or do not summarize it.
In Summarise Visually, a compatible standalone HTTP or HTTPS URL can follow a single-page article extraction route. That is not the same as crawling a website. For a product-focused overview, see the article summarizer for iPhone. This guide explains the broader source-checking method and the limits that determine whether one webpage can be summarized reliably.
Publication and evidence note: Heni Hazbay creates Summarise Visually and may benefit if readers download or subscribe. AI assistance supported Semrush research, source discovery, drafting, and editing. Product statements were checked against current first-party project evidence; general claims were checked against the cited official and research sources. Heni authorized publication on July 15, 2026. No current-device webpage walkthrough, extraction-coverage measurement, summary accuracy benchmark, or whole-site test was performed for this guide.
A webpage is not the same as a website
A website may contain thousands of pages, multiple authors, contradictory dates, private areas, interactive tools, databases, and continuously changing feeds. A single-page summarizer normally receives one URL or one block of text. Its result can describe only the material that was actually retrieved.
Define the scope before starting:
- Domain: example.org is the wider site and is too broad for a one-page summary.
- Page: example.org/reports/annual-review is one potential source.
- Visible article: the human-readable main text may be only part of that page.
- Extracted text: the summarizer may receive less or more than the visible article because of page structure.
- Summary: a generated interpretation of that extracted text, not a verified representation of the entire website.
Use precise language in your notes. Write “summary of the article at this URL, accessed July 15” rather than “summary of the website.” This avoids implying that navigation pages, linked evidence, updates, or other sections were reviewed.
Record the source before it changes
Web content can be edited, moved, or removed. Before generating anything, capture:
- the page title;
- author or responsible organization;
- publisher or domain;
- full canonical-looking URL shown in the browser;
- visible publication and update dates;
- your access date;
- the section or heading range you intend to summarize.
If the source is important, preserve a lawful reference copy. Apple documents ways to share a link in Safari and to save a webpage as a PDF. Saving a copy does not grant permission to republish it, and it may not capture every interactive element. It does give you a stable item to compare with the notes when the publisher permits that use.
Check whether the page is an original source, a report about another source, an opinion article, a product page, or an automatically assembled feed. If a page cites a study, law, filing, dataset, or announcement, open the linked primary material. A summary of a report about evidence is not a summary of the evidence itself.
Test whether the page is suitable
Before relying on URL extraction, answer these questions:
- Does the page load without a login, subscription wall, consent barrier, or challenge?
- Is the main content present in the initial HTML, or does it appear only after scripts run?
- Does the article have a clear title, author, date, and heading structure?
- Is the useful text on one page rather than behind tabs, an embedded viewer, or pagination?
- Are crucial facts contained in tables, charts, images, video, audio, or interactive controls?
- Are comments, recommendations, navigation, and advertising visually mixed with the article?
- Is the language supported by the intended workflow?
- Do you have permission to process the material?
If the answer reveals a material obstacle, do not assume the summarizer will solve it. Use accessible selected text you are entitled to process, a permitted PDF, or a manual reading workflow. Never bypass access controls.
How the verified website route works
Current project evidence supports a bounded article-URL process. For one standalone HTTP or HTTPS URL, the extractor fetches one HTML response and requires an HTTP 200 status. It prefers content inside an article or main element, removes selected page-chrome elements, gathers text, and keeps at most 15,000 characters.
Those details create important limits:
- a redirect, error, or blocked request can prevent extraction;
- a page rendered mainly by JavaScript may return little useful initial text;
- the preferred article or main container may be absent or incorrectly marked;
- removing page chrome can remove useful context;
- gathering paragraphs can omit tables, lists, captions, or footnotes;
- a long article can be truncated at the character boundary;
- only one response is fetched, so linked pages are not included.
The route is not a browser, JavaScript renderer, site crawler, paywall bypass, authenticated-content reader, or whole-domain summarizer. It should never be described as one.
The iOS share extension also has a selected-text route. When Safari supplies selected text, that selection can be transferred. Without a selection, the verified preprocessing fallback is only the page title and URL, not a guaranteed copy of the visible article body. If you want to summarize a specific passage, deliberately select the passage and confirm that the transferred text begins and ends where you expect.
A source-checked webpage workflow
1. State the question
Write what you need from the page: its main claim, a procedure, the result of an announcement, a policy change, or an overview of an argument. A defined question helps you notice when the output spends space on unrelated material.
2. Choose URL or selected text
Use the URL when the page is public, article-like, and represented in accessible HTML. Use selected text when you need one clearly bounded section or when automatic extraction includes clutter. Use a permitted PDF when stable pagination or a saved reference copy matters.
Do not combine several pages and describe the result as one source. If the article continues on another URL, record and summarize each page separately before creating a synthesis.
3. Inspect the received text
Compare the beginning, middle, and end with the webpage. Look for navigation labels, repeated footer text, missing lists, reversed columns, absent captions, and a cutoff before the conclusion. If you cannot inspect the received text directly, assume that extraction coverage is uncertain and keep the source open throughout review.
4. Preserve the page hierarchy
The W3C explains that headings communicate content organization and can support in-page navigation. Use the source’s headings as an audit outline:
| Source element | Summary check |
|---|---|
| Page title | Does the overview answer the same subject? |
| Introduction | Is the stated purpose preserved? |
| H2/H3 sections | Is each material section represented in the correct relationship? |
| Lists or steps | Are order, conditions, and exceptions retained? |
| Tables and figures | Were they checked manually rather than assumed from paragraph text? |
| Conclusion | Did a qualified conclusion become a stronger claim? |
Do not force a page with no meaningful headings into an invented hierarchy. Label the structure as your own if you create it.
5. Generate a first pass
Choose an output that matches the task: a short overview for triage, a detailed account for analysis, Key Points for a section map, or Q&A for candidate review prompts. Treat every output as a draft.
Keep source language around consequential terms. A privacy policy, medical explanation, financial statement, legal update, scientific result, or safety procedure can change meaning when a condition or exception disappears. For high-stakes use, a summary is not sufficient.
6. Verify claim by claim
Use the method in How to check an AI summary for accuracy. Split the summary into claims, then locate direct support on the page.
Check:
- names, roles, organizations, dates, and places;
- numbers, currencies, units, denominators, and time periods;
- quotations and whether they are verbatim;
- negation, uncertainty, and scope words such as “may,” “some,” and “only”;
- whether an association was rewritten as causation;
- whether the page reports someone else’s claim rather than endorsing it;
- whether a visible correction or update was omitted;
- whether linked evidence actually supports the page’s statement.
Research on factuality in abstractive summarization supports caution about fluent but unsupported generated content. The study does not measure Summarise Visually or guarantee that this checklist catches every error.
7. Add provenance and uncertainty
Finish with a source line that includes the title, author or publisher, URL, publication or update date, and access date. Add a note describing missing tables, unprocessed visuals, possible truncation, or sections you did not inspect.
If the page has changed since the saved or extracted version, do not silently merge them. State which version the summary represents and review the changes separately.
What to do when extraction fails
The page is dynamic
If the main text appears only after scripts load, copy a permitted passage or use an accessible publisher-provided print or PDF view. Do not claim the URL route processed material that was not present in the extracted response.
The page is paywalled or requires a login
Stop. Do not bypass the restriction or transmit private account content without authorization. If you lawfully have access and the publisher permits personal processing, work within those terms and the app’s supported input routes. A public summary should not reproduce protected material.
The page is too long
Divide it by genuine source headings. Record the range for each part and check whether the 15,000-character boundary excluded the ending. Summarize sections separately, then synthesize only after verifying each part.
The important content is visual
Review charts, diagrams, screenshots, maps, equations, and video separately. The verified article route is text extraction, not visual-page understanding. Add only observations you can support from the visual source and label them as your review.
The page has no clear author or date
Treat authority and freshness as uncertain. Search the site for an editorial or about page, but do not invent authorship. For time-sensitive claims, locate a dated primary source.
Rights, privacy, and responsible use
A webpage being readable does not mean its contents may be copied or redistributed without limit. The U.S. Copyright Office explains that fair use is evaluated case by case using multiple factors. Other countries apply different rules. Educational or personal intent is not an automatic permission, and this guide is not legal advice.
Avoid submitting confidential dashboards, private messages, patient information, client material, unpublished work, or authenticated pages unless you have authority and understand how the chosen service processes data. Remove secrets and personal data before using any external tool.
NIST’s Generative AI risk profile is general risk-management guidance, not a test of this product. Its relevance here is methodological: generated output should be governed according to the harm an error could cause. The higher the stakes, the closer the reader should stay to the original source and qualified expertise.
Limits and verification
The verified route handles one accessible HTTP or HTTPS page response. It does not crawl a domain, run arbitrary page scripts as a full browser, cross a paywall, access a logged-in account, follow linked sources, or guarantee that the main article was extracted completely. A 200 status means a response succeeded; it does not prove that the response contained every visible section in the right order.
The extractor can omit meaningful text or include irrelevant material. Long pages can be truncated at 15,000 characters. Tables, charts, equations, footnotes, captions, comments, and dynamic sections need separate review. Safari selected text can provide a deliberate passage, but without a selection the verified preprocessing fallback is only title and URL.
A source-checked summary can still be incomplete, biased, outdated, or unsuitable for a consequential decision. Check the live page, its update history, linked primary evidence, and any correction notice. No current-device page test or extraction benchmark was performed for this guide.
Final webpage-summary checklist
Before relying on the result, confirm:
- I summarized one identified page, not an implied entire website.
- I recorded the title, author or organization, URL, dates, and access date.
- I confirmed that I may process the material.
- I checked whether the page is dynamic, private, paywalled, paginated, or truncated.
- I compared the extracted beginning, middle, and ending with the source.
- I preserved the source heading structure where it was meaningful.
- I reviewed tables, figures, equations, footnotes, and visual evidence separately.
- I traced every important claim back to the page.
- I checked linked primary evidence instead of trusting a secondary description.
- I documented omissions, uncertainty, and the exact version summarized.
If any item is missing, label the result as a draft. A useful website summary is not merely short; it is bounded to one identifiable source and easy for another reader to verify.