Review / Market research / Permeate Lab
Found. Cited. Understood.
- Published
- September 15, 2026
- Revised
- September 16, 2026, version 1.0.1
- Format
- Review, Market research
Research review · Edition 1.0.1 · Source cutoff: September 15, 2026
This review evaluates research papers, including versioned arXiv manuscripts, and selected provider guidance for people deciding how to make business evidence discoverable and useful to humans and AI systems. It contains no original Permeate field measurements and makes no claim of measured visibility gains. The title's “understood” refers to observable fidelity, not hidden model cognition.
A company can publish accurate research, make the page accessible, and still never appear in an AI answer. It can also appear in a citation while the answer loses the qualification that made its evidence useful.
Those outcomes need different diagnoses. Our reading of the evidence supports a practical proposition: measure access, discovery, citation and claim support separately before deciding what to change. It does not support a universal recipe for gaining AI visibility.
Recent research also suggests that helping one stage can damage another.
A citation is one event in a longer process
| Observable question | Evidence to retain | Limit of the observation |
|---|---|---|
| Can the public page be read? | Response, readable text and time of request | Our successful fetch does not prove a provider ingested it. |
| Did the system encounter the source? | Visible retrieval result or trace, if available | A proprietary system may expose only part of its process; a missing trace does not establish that the source was not encountered. |
| Did the answer cite the edition? | Answer, citation destination and resolved version | A reference can be present without supporting the attached claim. |
| Did the answer preserve the evidence? | Claim-level comparison, including scope and qualifications | Supported wording does not reveal the model's internal causal dependence. |
| Did the result help someone? | Observed task completion, referral or other defined outcome | More citations alone do not establish usefulness or business value. |
This practical framework is not an original academic model or authority score. Earlier research separates citation support from reference presence; its historical results do not rate September 2026 products. Evaluating Verifiability in Generative Search Engines.
What changes when retrieval is included
The original GEO study found that changes to source presentation improved its visibility measures under tested conditions. It included a historical Perplexity evaluation, but did not evaluate effects on search rankings. That distinction limits what a publisher can infer: improving how an already selected document is represented is not the same experiment as helping a new page become discoverable. GEO, methods and limitations.
SAGEO Arena tests retrieval, reranking and generation together. Its August revision reports that body-only optimization often damages upstream visibility across BM25, dense and hybrid retrieval configurations. That is broader than a keyword-matching result, but the experiment still selects targets from documents already reaching its baseline generation stage and does not reproduce proprietary commercial ranking signals. It supports testing the whole path, not promising that a new page will enter it. SAGEO Arena v2, §5.4 and Appendices B–D.
Our interpretation is to preserve useful specificity and test the failure we can observe. Rewriting everything into a supposedly ideal AI style would make it difficult to know which change helped, which hurt, or whether any result came from the edit at all.
AgentGEO points toward localized diagnosis and repair. Its evaluation uses a controlled simulation, and the paper leaves commercial-engine validation to future work. We take its held-out query approach as a design lesson; we do not take its reported improvement as an expected Permeate result. AgentGEO, section 6.1 and conclusion.
These studies are not a vote with a common denominator. They test different systems, outcomes and intervention conditions. Their useful convergence is the need to diagnose stages. Their material divergence is whether a particular edit survives a different retrieval and generation environment.
What different experiments can establish
Large experiments can answer narrow questions. A 252,000-trial study measures which of exactly two supplied documents receives the first citation, without running search. Repeated trials within model configurations are not independent systems. Its paired design examines first-citation preference; it cannot establish how a publisher enters the candidate pool. We use it as a design lesson, not a universal optimization checklist. What Gets Cited, §2.2 and §5.
Production evidence also exists. Pinterest reports traffic gains from annotation and interlinking changes on its own visual platform. That is relevant operational evidence, but its scale, bundled intervention and incomplete public experimental detail prevent us from adopting its reported lift as a forecast for Permeate. Pinterest framework, §4.3.
| Evidence type | Useful conclusion | Boundary |
|---|---|---|
| Two-document first-citation experiment | An edit changes first-citation preference within its tested pair | Does not establish web discovery or all citation use |
| Controlled retrieval pipeline | Stages can respond differently to the same edit | Does not reproduce every commercial ranking signal |
| Deployed-system observation | A recorded product configuration behaved a certain way | Does not isolate the cause of a publisher's visibility |
| Publisher intervention report | A reported operational change merits investigation | Effect transfer requires comparable context and inspectable methods |
These are different units of evidence. Revisions of one study, reposted papers and summaries of the same experiment remain one study family.
Being used accurately is harder to establish than being linked
The Tow Center's 2025 investigation found attribution failures in a defined news-excerpt retrieval task. Its task design is useful for testing known-source identification, but it does not estimate error across every kind of AI question today. Tow Center methodology and findings.
A July 2026 study of deployed search systems adds a temporal warning: similar topic coverage can coexist with changing sources. Its observations were collected in 2025 through specified search and API configurations, so they do not rate today's consumer interfaces. For anyone measuring current performance, the implication is to retain every attempted query, distinguish available search from observed search, and report variation across repetitions. Characterizing Web Search, methods and §6.
Three outcomes should remain distinct:
- Supported: the cited source provides evidence for the claim as stated.
- Qualified: the answer retains the scope, date and uncertainty needed to interpret that claim.
- Causally attributable: the source actually caused the answer's content. We generally cannot establish this from a public response alone.
An answer can satisfy the first two without allowing us to establish the third. Fidelity also matters to a human reader: preserving a number while losing its population, timeframe or uncertainty changes what the evidence says.
Visibility research also examines influence
September research broadens the question from gaining visibility to evaluating influence. Counter-GEO-Bench examines information-distorting rewrites in a controlled, English-language benchmark. It explicitly excludes end-to-end commercial products, so its results cannot establish how a named consumer service currently defends itself. Counter-GEO-Bench, limitations.
For Permeate, the implication is an editorial constraint: preserve the evidence and disclose its provenance even when a more persuasive version might be easier to cite. Citation gain accompanied by a missing limitation is a defect in faithful representation. Reposting our material on another property will remain owned distribution, not independent corroboration.
That constraint also changes the business interpretation. An assistant that answers well by citing the original provider documentation may serve its user better than one that cites our summary. Permeate's absence can be commercially disappointing without being an error in the answer. Those two outcomes deserve separate reporting.
What publishers can verify now
Google's current guidance keeps ordinary SEO relevant and does not require a special writing format or llms.txt for its generative Search features. Google states that llms.txt does not improve visibility or ranking in Google Search; this claim does not extend to every assistant. For Google, check that the page is indexed and snippet-eligible and that the site is included in Search generative AI features in Search Console. These Google-specific controls do not establish a universal formatting requirement. Google's optimization guidance.
Provider controls and measurements also have different scopes:
| Surface | Verified distinction | Measurement consequence |
|---|---|---|
| Google Search | Its report measures link impressions in AI Overviews and AI Mode, not direct counts of citation events or referrals. Unavailable values can become zeros in exports. | Retain report availability and UI evidence alongside downloaded numbers; keep Gemini observations separate. |
| ChatGPT | OAI-SearchBot supports search; ChatGPT-User handles user-requested actions; GPTBot covers potential training. | A user-triggered fetch or permitted crawler is not evidence of ordinary selection and guarantees neither indexing nor citation. |
| xAI tools | Encountered source URLs differ from final-answer citations. | Count the citation actually attached to the answer; API evidence remains separate from consumer Grok observations. |
| Connected research tools | Lovable documents publishing app tools through MCP. | Deliberate access tests whether the connection works. It does not demonstrate unprompted discovery. |
Sources: Google report, OpenAI crawlers, xAI citations, Lovable MCP integration. These are documentation checks, not proof that Permeate's accounts or integrations are configured. A documented feature is not evidence of a working connection or an observed research outcome.
A practical decision sequence
Start with a useful question, bounded answer, inspectable sources and visible date. Resolve factual and scope errors before optimizing presentation.
Diagnose the observable stage. Successful supplied-URL reading without ordinary discovery raises an exposure or retrieval question, not a proven cause. A citation with an unsupported claim raises a fidelity question. Accurate answers using original documentation may need no correction.
Keep changes inspectable. Publishing an article while changing schema, creating distribution copies and activating a research connector bundles several interventions. A subsequent change in citations cannot isolate which intervention mattered. Record what changed and when, retain adverse results, and compare like conditions before making performance claims.
Distribution should make the evidence easier to inspect. Put the primary source beside the claim it supports. Give an archive, dataset or technical companion a distinct purpose and an explicit relationship to the original edition. Copies uploaded by the same publisher remain one evidence family. Additional domains do not make them independent corroboration.
Each derivative also needs an edition and correction history. Preserving a stale claim across more surfaces undermines useful access.
These recommendations follow the reviewed evidence and its limitations. Permeate has not experimentally tested their combined effect. An original study would need separate questions, conditions, analysis and results.
How this review was conducted
The unit of analysis is a study, versioned research report or provider statement, not a company, user or market population. We selected English-language original research and official documentation relevant to discovery, citation and faithful reuse, with source grounding through September 15, 2026. The cited research spans 2023–2026; publication dates and observation dates are distinguished wherever they materially affect interpretation.
This was a targeted critical review. Search and citation-following deliberately sought contrary evidence, including retrieval degradation and influence-defense research. Original methods, results and limitations were inspected where available. Provider documentation was used for the provider's own stated controls, not as experimental proof of performance. Secondary surveys helped locate originals and were not counted as additional experiments.
Searches combined generative-engine optimization, citation faithfulness, retrieval, benchmarks and live-search experiments, with arXiv and ACL discovery queries. Citation-following and a direct AgentGEO identifier lookup supplemented them. The exact identifier-query syntax was not retained in the research summary, so the search history is incomplete. The review was not preregistered; its decision scope and reporting boundaries were clarified during synthesis rather than locked before initial searches.
Inclusion required an identifiable original task or an official statement about the provider's own system. Vendor playbooks, unsupported uplift promises, snippets without readable originals and duplicated study versions were excluded from substantive effect claims. The extension stopped after inspecting contrary retrieval evidence, a deployed-system study, a controlled citation experiment and a production publisher report. This was a practical coverage decision, not demonstrated exhaustion of the literature.
The synthesis compares study designs and their boundaries rather than pooling incompatible outcomes. It is not a systematic review or meta-analysis. There was no exhaustive database export, independent duplicate screening, benchmark replication or original consumer-product data collection. Selection and publication bias remain possible; English-language selection and opaque commercial systems constrain the conclusions. No numerical certainty score or representative market estimate is assigned.
Ownership, AI use and corrections
Permeate Lab, a Nova3 AI research practice, commissioned and publishes this review. Permeate has a commercial interest in research discoverability, which makes both adverse evidence and the distinction between distribution and independent corroboration material. No client outcomes or private client data are used.
AI agents assisted source discovery, synthesis, drafting, source checks and editorial review. Mark Jones, Permeate's owner, authorized publication through a delegated AI editorial workflow. Separate agent checks are not independent human substantive review or academic peer review; publication authorization does not mean Mark performed a substantive review of the final text. The review does not claim external researcher endorsement, certification or validation of a completed field study. Its linked originals let readers inspect the evidence supporting its central factual claims.
For a correction, use the Permeate inquiry form and identify this title, edition, disputed passage and supporting source. Material corrections should identify the affected claim, explain what changed, and date the replacement edition while preserving the prior record. The source cutoff is a freshness boundary, not a promise that provider behavior or documentation remains unchanged.
Change note
Editorial update: standardized the Nova3 AI affiliation and apostrophe styling. Research findings and source links are unchanged. Recorded September 16, 2026.
Publisher: Permeate Lab, a Nova3 AI research practice. Format: Review. Version 1.0.1. Published September 15, 2026.