Why your technical docs may be your most valuable digital asset

Documentation answers the questions buyers actually ask, which makes it bottom-of-funnel content. What LLM traffic is worth, what retrieval rewards, and which pages to fix first.

GTM
August 18, 2026
·
6 min
read
Abstract brand form on a purple field

Your documentation is bottom-of-funnel content, and most companies still fund it like a cost center they'd close tomorrow if the support queue allowed it. It gets a contractor, a backlog, and apologies. Meanwhile the buyer three weeks into evaluating you is reading it at eleven at night, working out whether the product does the specific thing they need, and they'll never fill in a form to ask.

That was true long before generative engines existed, and what changed is that a second reader showed up, reading the same pages for the same reason in roughly the same way. Retrieval doesn't read a website the way a person reads a brochure, front to back and in the order you chose. It works more like someone pulling one part off a shelf in a workshop, where the labelled shelf wins and the unsorted pile loses, however good the parts in that pile happen to be. Documentation is the labelled shelf. Most of the rest of a B2B website is the pile in the corner.

So what's that traffic actually worth, and which pages do we fix first?

Is small traffic bad traffic?

Not when it converts. Patrick Stox published Ahrefs' own analytics for a 30-day window in June 2025, and AI search sent 0.5% of their total traffic while producing 12.1% of their signups, a conversion rate 23 times higher than traditional organic search. Twelve percent of signups from half a percent of visits!

Should we rebuild a website around half a percent of its traffic? Well, of course not, and Stox himself states two caveats that are worth reading before anything gets moved. Visitors from AI search click links about 75% less than they do in organic search, and he isn't sure the pattern scales, quoting a colleague who expects these click rates are the highest they'll ever be. So the multiple will come down. The direction holds anyway, and it matches practice: this traffic is small in volume, late in the cycle, and it lands after the comparison work has happened somewhere we can't watch.

Which is why the number to manage is qualified traffic rather than raw traffic, tagged with UTMs and judged on the leads it converts. This reader turns up at the end of their research rather than at the start of it. So the pages worth optimizing are the end-of-funnel ones: pricing, comparisons, integration limits, security posture, and the documentation that answers what a sales call would have answered.

What makes a document trustworthy is what makes it retrievable

An engineer scanning a reference page wants one section, wants to confirm it's current, and wants to leave. They skim headings. They trust a definition that defines, they distrust a page that buries the specific inside the general, and they abandon anything that reads like it was written to fill a content calendar. Retrieval behaves in a recognizably similar way, because it works with pieces of your content rather than the whole polished narrative, and the pieces it can use are the ones that stand on their own.

So why do the two disciplines keep issuing the same instructions? Because they're solving the same problem: a reader who needs a precise answer and has no patience for the paragraph it's buried in. No fluff, high accuracy, strong structure, lean content, clear hierarchy.

What a technical reader needsWhat retrieval needsThe move for both
To land on one sectionPassages that stand aloneH2s phrased as the question
Terms defined the same way everywhereOne canonical passage per conceptA definitions page everything links to
Proof the page is currentFreshness it can checkLast-updated dates, versioned pages
Values, not proseStructure it can liftTables for limits, tiers, specs, versions
The next answer one click awayLinks between related answersDeliberate internal linking

I've built that shape several times, and none of it was built for retrieval, because retrieval wasn't a consideration in most of those years. At Lightrun, a 250-page knowledge base from a blank slate in eight weeks. At Firebolt, 100+ pages standing before public launch, and at Lytx and Surfsight, a 300+ page developer portal. At Coro, a compliance hub covering more than 20 global regulatory frameworks, built as an always-on resource rather than a campaign. At Visual Layer I documented the entire product codebase, and it came back from the engineers who wrote that code with close to zero corrections. That's the only quality test for technical content that's ever meant much to me.

None of this is a new argument either. A deck I wrote in March 2021 had a slide titled "Doc use cases: the bottom line", and its first line read "leverage docs for additional SEO + POCs and decision-making". So were docs funded as a discovery channel back then? They weren't, and mostly they still aren't, even though they were already the layer that buyers evaluated us on.

So is llms.txt the fix?

Not on its own. A mapping file or an llms.txt is an internal planning tool that forces us to decide which page owns which question. That's genuinely useful, and it's mostly useful to us rather than to the machine, because the work that changes outcomes is duller.

  • FAQs written the way a person asks. Somebody types "does it work behind a proxy", not "network configuration considerations".
  • The follow-up question, answered underneath it. Every answer produces a second one.
  • Target terms inside the H2s. Structure carries as much weight as keyword presence.
  • Internal links between the answers. Cheap, unglamorous, and usually skipped.
  • Natural writing instead of stuffing. A relevant subject contains the right terms anyway.

One warning belongs in bold here: don't damage your Google rankings while you optimize for answer engines. SEO isn't over. It has been declared dead several times before, and every time the answer came back the same, which is that it needs continuous adjustment rather than a funeral.

How do we know any of it is working?

Partly through tooling, carefully: on a competitive engagement for a B2B SaaS client we mapped their visibility against five competitors and evaluated five AI visibility and brand-monitoring tools along the way. The tools disagree with each other, sometimes badly, which is what you'd expect from a category this young. So we treated their numbers as a baseline to move rather than a score to win, and put the weight on GA4 and UTM-tagged traffic, where conversions are countable and arguments are shorter.

The limit is worth stating plainly, because restructuring your documentation won't rescue positioning that was never clear in the first place. It won't make a model recommend you over a genuinely better product, and it can't repair a claim your buyers already doubt. What it does is make you findable, quotable, and easy for a skeptical reader to verify. Whether the reader likes what they verify is a separate piece of work, and it happens well upstream of anything in this post.

Key takeaways

  • Docs answer evaluation questions, which makes them bottom-of-funnel content.
  • Qualified traffic is the number that matters, so tag it with UTMs and count leads.
  • Technical writing and answer-engine optimization ask for the same things: accuracy, hierarchy, definitions, tables, links.
  • llms.txt is planning. FAQs, follow-up questions, headings, and internal linking are the work.

The Snyk post on PCI and open source security requirements was built this way in 2019, long before anyone said GEO, and it holds up because the structure did the work rather than the adjectives. If your documentation is accurate and nobody can find anything in it, that's fixable, and it's cheaper than the campaign you were about to run instead. Bring the sitemap to a free 30-minute call, and we'll pick the ten pages your buyers evaluate you on and decide which question each one owns.

Mentoring, free forever.

I mentor marketers and technical writers. Breaking in, growing your career, automating with AI and more.

Book a free session
zero-shot-learning
Zero-shot Learning
A machine learning setup where a model attempts to classify objects or perform tasks it has never explicitly seen during training, often by using semantic descriptions or knowledge transfer.
Text Link
zero-defect-manufacturing-zdm
Zero Defect Manufacturing (ZDM)
A manufacturing philosophy aiming for zero defects through continuous monitoring and data-driven process improvement.
Text Link
workflow-automation
Workflow automation
The use of technology to automatically perform tasks that were previously done manually, such as data triage.
Text Link
warm-cache-vs-cold-start
Warm Cache vs. Cold Start
A warm cache means data is already loaded into memory, while a cold start means the system must fetch data from disk.
Text Link
visual-search
Visual Search
The process of finding similar images and/or clusters for a given query image or cluster.
Text Link
video
Video
A sequence of images that enables temporal analysis, motion detection, and event recognition across time.
Text Link
vector-databases
Vector Databases
Systems that use vector infrastructure where embeddings are created and stored, often within a vendor's backend system.
Text Link
unsupervised-learning
Unsupervised Learning
Machine learning where the model identifies patterns and structures in data without the use of labeled examples or ground truth.
Text Link
unindexed
Unindexed
When data is not organized in a way that makes it easily searchable or retrievable.
Text Link
transfer-learning
Transfer Learning
A technique where a model trained for one task is reused as the starting point for a model on a different but related task.
Text Link
training-signals-feedback-loss
Training signals (Feedback / Loss)
The information or feedback a model learns from during training. For a supervised learning model, the training signal is the ground truth label.
Text Link
throughput
Throughput
The number of operations or data items a system can process per unit of time.
Text Link
threshold
Threshold
A similarity value over which we define images as similar, identical, or close-enough.
Text Link
technical-debt
Technical debt
The accumulated cost of shortcuts and workarounds (like homegrown scripts) that make a system difficult to maintain over time.
Text Link
synthetic-data
Synthetic Data
Artificially generated data created by algorithms rather than collected from the real world.
Text Link
supervised-learning
Supervised Learning
A method of machine learning where models are trained using labeled data, meaning the correct answers (ground truth) are known. This is fundamental to most CV tasks.
Text Link
structured-datasets
Structured datasets
A dataset organized in a clear, consistent way with well-defined fields and structure.
Text Link
sql-structured-query-language
SQL (Structured Query Language)
A standardized programming language used for managing and manipulating relational databases.
Text Link
similars
Similars
Images and clusters that are visually similar to a chosen vertex (query item).
Text Link
similarity-cluster
Similarity Cluster
A cluster of data entities grouped by their similarity.
Text Link
similarity
Similarity
How close two images are to one another in terms of appearance and content, usually a value between 0-1 calculated between embeddings.
Text Link
siloed-workflows
Siloed workflows
When different teams work in isolation using separate tools that don't connect, causing friction and delays.
Text Link
semantic-segmentation
Semantic Segmentation
Classifying every pixel in an image into a category, creating a detailed pixel-level understanding.
Text Link
semantic-search-nlp-search
Semantic search (NLP Search)
A search method that finds results based on the meaning or intent of a query, not just on matching keywords.
Text Link
self-supervised-learning
Self-Supervised Learning
Learning approach enabling a model to generate its own supervisory signals from unlabeled data by predicting specific parts of an input based on other parts.
Text Link
selected-item
Selected item
The item/s (image, cluster, or object) that a user selects to see details about or perform an action on.
Text Link
seed-anchor-prototype
Seed (Anchor / Prototype)
In active learning, a seed is an initial labeled example manually selected for a class. Seeds are the anchors a system uses to recognize patterns and start label propagation.
Text Link
scalability
Scalability
How big a model is and how efficiently it can run, involving parameter count and memory footprint.
Text Link
s3-simple-storage-service
S3 (Simple Storage Service)
AWS's object storage service used for storing and retrieving data at scale.
Text Link
roi-return-on-investment
ROI (Return on Investment)
A measure of how much profit or benefit is gained compared to the cost invested.
Text Link
remediation
Remediation
The process of fixing problems or correcting errors in data, such as re-labeling or removing bad samples.
Text Link
reinforcement-learning
Reinforcement Learning
Learning through trial and error by receiving rewards or penalties for actions taken in an environment. In vision AI, it's used when systems must learn sequential decision-making.
Text Link
reinforcement-learning-rl
Reinforcement Learning (RL)
A machine learning paradigm where an agent learns to take actions in an environment to maximize cumulative reward. Learning happens via trial-and-error feedback rather than labeled examples.
Text Link
regression-model
Regression Model
A model that predicts a continuous value instead of a category. It answers "how much?" or "how many?"
Text Link
recall
Recall
The ratio of correctly predicted positive observations to the all observations in the actual class. Also called Sensitivity. It measures how many actual positives were captured.
Text Link
ranking-model
Ranking Model
A model that orders a set of items by relevance or importance. It decides what should come first, second, third.
Text Link
ram-random-access-memory
RAM (Random Access Memory)
A form of computer memory that can be read and changed in any order, used to store working data and machine code currently in use.
Text Link
query-item-vertex
Query item (Vertex)
The item (image, cluster, or object) that a user chooses to see its similar images.
Text Link
production-models
Production models
An AI model that is ready for use in a real-world application, as opposed to one that is still in development.
Text Link
privacy-safeguards
Privacy safeguards
Privacy safeguards ensure sensitive personal information (PII) is anonymized or replaced before training.
Text Link
predictive-ai
Predictive AI
A type of AI that uses statistical algorithms and machine learning techniques to analyze historical data and make predictions about future outcomes.
Text Link
prediction
Prediction
An attempt by a model to replicate the ground truth, usually accompanied by a confidence score.
Text Link
precision
Precision
The ratio of correctly predicted positive observations to the total predicted positives. It measures how accurate the positive predictions are.
Text Link
postgresql-postgres
PostgreSQL (Postgres)
A powerful open-source object-relational database system (OLTP) optimized for transactional reliability.
Text Link
polygon
Polygon
A (usually non-rectangular) region defining an object with more detail than a rectangular bounding box.
Text Link
pipeline
Pipeline
The end-to-end process of going from raw images to a prediction, including collection, annotation, training, and deployment.
Text Link
pg-postgresql
PG (PostgreSQL)
Common abbreviation for PostgreSQL, an advanced open-source object-relational database system known for reliability and feature robustness.
Text Link
petabytes-pb
Petabytes (PB)
An extremely large unit of digital data (one million gigabytes). Modern visual datasets are often measured at this scale.
Text Link
perceptual-duplicates
Perceptual duplicates
Images that are not identical but are visually very similar, often the same photo resized, compressed, or color-altered.
Text Link
p50-p95-50th-95th-percentile
p50 / p95 (50th / 95th Percentile)
Statistical metrics used to measure system latency or performance. p50 represents the median (typical) performance, while p95 represents the "tail" latency (worst-case for 95% of requests).
Text Link
overfitting
Overfitting
When a model learns the training data (including noise/errors) so well that it fails to generalize to new, unseen data.
Text Link
outliers
Outliers
Images that "don't belong" to the rest of the data, such as corrupted files or extreme edge cases.
Text Link
online-transaction-processing-oltp
Online Transaction Processing (OLTP)
Database systems optimized for managing short, atomic transactions like inserts, updates, and deletes.
Text Link
online-analytical-processing-olap
Online Analytical Processing (OLAP)
Database systems optimized for reading and aggregating large datasets, designed for analytical queries.
Text Link
on-prem-on-premises
On-Prem (On-Premises)
Deploying software on local servers or private infrastructure rather than the public cloud, often for security or latency reasons.
Text Link
object-detection
Object Detection
A core computer vision task that combines classification (what is it?) and localization (where is it?) to identify and precisely locate objects within an image, typically using bounding boxes.
Text Link
object
Object
A region of interest in an image, usually containing an instance of a specific class.
Text Link
noisy-data
Noisy data
Data that is corrupt, irrelevant, or contains errors that can confuse a model.
Text Link
neural-networks
Neural Networks
Models composed of interconnected layers of weighted computations (“neurons”) that learn nonlinear mappings from inputs to outputs. Networks can be shallow or deep; depth enables learning increasingly abstract features.
Text Link
neural-network-ann
Neural Network (ANN)
Artificial neural networks are a subset of machine learning inspired by the human brain, mimicking the way biological neurons signal to one another.
Text Link
natural-language-processing-nlp
Natural Language Processing (NLP)
The area of AI focused on processing, understanding, and generating human language (text and sometimes speech). Includes tasks like classification, extraction, summarization, translation, and question answering.
Text Link
multimodal-ai
Multimodal AI
Systems defined by their ability to process and integrate diverse data types simultaneously, such as vision, language, and audio.
Text Link
modeling-approaches-paradigm
Modeling Approaches (Paradigm)
Different modeling approaches describe the ways algorithms are designed to process data and learn patterns.
Text Link
model-types-architecture
Model Types (Architecture)
In machine learning, model types define the general way an algorithm learns from data and makes predictions (categorizing, generating, predicting values, or ranking).
Text Link
model-drift
Model Drift
The degradation of a model's performance over time as the environment or data changes.
Text Link
mislabeled
Mislabeled
When a piece of data is incorrectly tagged or categorized. This is a severe issue as it teaches the model incorrect patterns.
Text Link
metadata-meta-info
Metadata (Meta-info)
Descriptive information about data, such as production parameters, timestamps and file locations, often stored alongside images in CSV or JSON files.
Text Link
machine-learning-operations-mlops
Machine Learning Operations (MLOps)
A set of practices for reliably and efficiently deploying and maintaining machine learning models in a production environment.
Text Link
machine-learning-ml-2
Machine Learning (ML)
An AI technique that teaches computers to learn from experience (data) without being explicitly programmed.
Text Link
machine-learning-ml
Machine Learning (ML)
A field of AI where models learn patterns from data to make predictions or decisions without being explicitly programmed for every rule.
Text Link
localization
Localization
Identifying where in an image an object resides, providing x/y coordinates (as opposed to just the class label).
Text Link
layer
Layer
Made up of neurons (or blocks of neurons) that compose deep neural networks. Adding layers makes a network "deeper" and increases predictive power.
Text Link
latency
Latency
The time delay between input and output—how long a system takes to process a request.
Text Link
large-language-model-llm
Large Language Model (LLM)
A very large AI model designed to understand and generate human language, trained on vast amounts of text data. In computer vision, they are used in multimodal systems for captioning or visual QA.
Text Link
label-propagation-lp-semi-supervised
Label Propagation (LP / Semi-Supervised)
A semi-supervised learning technique in which a system learns from a small set of labeled seed examples, then assigns labels to similar unlabeled data based on confidence scores and improves through feedback loops.
Text Link
label-tag-class
Label (Tag / Class)
The specific category assigned to a single image or data point. A label is the granular application of a class to one item.
Text Link
k-nearest-neighbors-knn
k Nearest-Neighbors (kNN)
The number (k) of images that are the most similar to the image in question.
Text Link
job-production-metadata-job-data-header-info
Job/Production Metadata (Job Data / Header Info)
In manufacturing inspection data, these are the configuration parameters (Job, Setup, Recipe) that identify where images came from and keep a single dataset consistent.
Text Link
iteration-cycle-loop
Iteration (Cycle / Loop)
In active learning and label propagation workflows, an iteration is one complete cycle: seed selection → automated labeling → human review → refinement.
Text Link
issue-type
Issue Type
A specific type of problem that generates zero or more issue instances during the analysis process.
Text Link
issue
Issue
A concrete instance of an issue type, associated with one or more images, allowing teams to prioritize remediation.
Text Link
instance-segmentation
Instance Segmentation
Identifying individual object instances and their pixel-level boundaries separately, rather than just classifying pixels into categories. It combines object detection and semantic segmentation.
Text Link
inference
Inference
The process of applying a trained model over a dataset to make predictions on new, unseen data.
Text Link
image-recognition
Image Recognition
The ability of a computer vision system to identify what an image contains, such as objects, people, places or text, and assign it one or more labels.
Text Link
image-classification
Image Classification
See Image Recognition.
Text Link
image-augmentation
Image Augmentation
A technique that creates new training data by making minor alterations to existing images to make the model more robust.
Text Link
image-attributes
Image attributes
Simple features based on basic math or rule-based calculations for measuring characteristics like brightness, blurriness, or contrast.
Text Link
image
Image
A file that contains a visual representation of something; the fundamental input data for computer vision.
Text Link
iiot-industrial-internet-of-things
IIoT (Industrial Internet of Things)
A network of connected industrial devices that collect, share, and analyze data to optimize operations.
Text Link
human-in-the-loop-hitl
Human-in-the-Loop (HITL)
A workflow that combines human expertise with machine automation, often for reviewing uncertain predictions.
Text Link
homegrown-tools
Homegrown tools
Custom software or scripts built by a company's internal team to solve a specific problem. They often become fragile and hard to maintain.
Text Link
holdout-sets
Holdout sets
A simple method where part of the dataset is set aside as unseen data to test the model.
Text Link
heuristics
Heuristics
Rule-of-thumb approaches that rely on handcrafted logic rather than learned patterns.
Text Link
ground-truth-gt-gold-standard
Ground Truth (GT / Gold Standard)
The verified, correct labels for a dataset, usually confirmed by domain experts. It is the authoritative reference that models are trained on and that predictions are measured against.
Text Link
generative-model
Generative Model
A model that creates new data resembling the data it was trained on. It doesn't just sort or predict—it "imagines" new examples.
Text Link
generative-adversarial-network-gan
Generative Adversarial Network (GAN)
A type of generative model with two competing parts: a "generator" that creates new data and a "discriminator" that tries to distinguish real data from fake.
Text Link
full-text-search-fts
Full-Text Search (FTS)
A database capability that allows text queries over large documents or metadata fields using tokenized indexes.
Text Link
foundation-model
Foundation model
A very large, general-purpose model trained on massive amounts of data, later adapted to specific tasks.
Text Link
fine-tuning
Fine-tuning
Fine-tuning involves taking a pre-trained model and continuing its training on a smaller, specific dataset for a specialized task.
Text Link
filtering
Filtering
Process of reduction of a set of items by applying conditions (query expression) to items' metadata properties.
Text Link
Keep reading

Tell me what you're building and what's in the way. I read everything.