Levertrace Lab

← Back to the index

Case 15 · Direction I · Owned-page & structured levers · No visible movement

How can a business test one optimisation change

A useful LLM optimisation test starts with a recorded before state, changes one public lever, repeats named prompt and language conditions, compares sources, and labels uncertainty instead of forcing a clean verdict.

Recorded by Salomé Rivecourt June 9, 2026

The cleanest LLM optimisation test is often small enough to feel unsatisfying. One change, one before state, one after run, and a careful refusal to explain more than the evidence allows.

A French local service company rewrites its homepage on Monday. On Tuesday, it corrects a directory profile. On Wednesday, it adds structured data. By Friday, someone asks an LLM what the company does and gets a better answer. The team wants to celebrate. Levertrace Lab would slow the room down, not because the movement is meaningless, but because the test has already become tangled.

Another composite case begins more plainly. The business changes one category sentence on a service page and leaves the directory alone. Before the change, the model calls it a general maintenance provider. After the change, under the same French prompt, one answer uses the narrower category. The English answer stays vague. This is not a grand result. It is something better for research: a small piece of movement with a visible edge.

Start by recording the before state

The before state is the part most businesses skip. They remember that the old answer was wrong, but they do not preserve the wording, prompt, language, date or source condition. Later, when the answer shifts, nobody can say whether the model changed, the prompt changed, or the memory of the old answer softened in the retelling.

Levertrace records the before state as an observation. An observation may be a full answer, a source mention, a phrase choice, an omission, a refusal or a mismatch between French and English wording. The lab keeps the raw wording because small phrasing differences matter. “General maintenance,” “building services” and “lift-door maintenance” are not interchangeable labels for a company trying to be understood.

A practical before state contains the prompt family, the exact or near-exact prompt, the language condition, the visible source condition and the date of observation. If the interface cites sources, those sources are logged. If the interface does not show retrieval, the answer is marked source-uncertain. The lab avoids guessing. A confident answer without citations is still an answer without visible citations.

For a French business, the before state should usually include both French and English prompts when language transfer matters. The French answer may already be correct while the English answer is stale. Or the English sales page may dominate the answer while French documentation carries the more precise category. Without recording both, the business may fix the wrong surface.

A before state — in Levertrace’s working definition — is the recorded answer and source context captured before a single public lever changes, because later movement has no meaning without a baseline. It is not a memory of what someone thought the model used to say.

Change one lever, even when several are tempting

A lever is the public optimisation change being tested. It may be an owned-page edit, a directory correction, a structured data change, a bilingual rewrite, an interlinking change, a dated update or a repeated claim. In ordinary marketing work, changing several at once may be sensible. In causal observation, it makes the finding muddy.

The discipline feels artificial because real websites are messy. A founder sees a wrong description and wants to fix the homepage, rewrite the about page, update schema, correct directory listings and ask a partner to change old wording. That may be the right operational response. It is not a clean test of one lever. Levertrace separates these activities when the goal is to learn what moved the answer.

Study object A, a composite French local service company in building maintenance, shows the problem. Suppose the company corrects its official service page and a local directory in the same week. An LLM answer later adopts the new category. Did the owned page move it? Did the directory move it? Did the agreement between the two matter? The lab can record an observed shift, but attribution becomes weak.

A cleaner version changes only the directory listing first. The website stays the same for the test window. If a live-search answer then adopts the new category while a non-browsing answer does not, the lab can label the movement as possible directory influence under a live source condition. That is still cautious. It is also much more useful than a broad claim that “AI visibility improved.”

Study object B, a composite French B2B software company, gives a different version of the same lesson. If the company updates English landing pages, French documentation and structured data together, any later answer change is difficult to read. If it first changes one English positioning sentence, the lab can observe whether that sentence affects English answers, French answers or neither.

The point is not to freeze all business work for the sake of research. The point is to decide when the business is testing and when it is simply fixing. Those are different modes.

Repeat the prompt without worshipping identical wording

Repeatability in Levertrace work means another reader can understand the comparison setup. It does not mean every model must return identical wording. LLM answers are too unstable for that, and a method built on pretending otherwise will break the moment a second run gives a slightly different sentence.

The lab uses prompt families. A prompt family might include “What does this company do?”, “Describe this business in one paragraph,” and a French equivalent asking for the activity and service category. The prompts are close enough to test the same business description, but varied enough to catch brittle answer movement. If only one oddly phrased prompt shows the correction, the lab treats the result carefully.

Source condition is recorded for each run. Live-search behaviour is separated from non-browsing or source-uncertain behaviour whenever the interface makes that distinction visible. A corrected source may appear quickly in live search while memory-shaped answers stay old. That is not failure or success by itself. It is a split observation.

Language condition is also named. French prompt, French source, French answer is one condition. English prompt, French source, English answer is another. English prompt, English source, English answer is another again. This may sound fussy until the business sees a correct French answer and a crooked English summary in the same hour.

The lab also watches omissions. If a model avoids the company after the change, the absence is logged. If it names the business but refuses to describe it specifically, that is logged too. Silence can be as informative as a wrong category when the question is whether the business has become legible to the model.

Compare the after state against the source trail

The after run is where many teams rush. They look for the new phrase and stop reading once they find it. Levertrace reads the whole answer, then compares it with the surrounding public sources. A model can include the new wording and still carry an old source conflict in the same paragraph.

The lab asks which fact moved. Category? City? Service boundary? Customer type? Ownership? Language? A business description can improve in one part and remain stale in another. If the tested lever was a category edit, a city correction in the answer may be incidental. If the tested lever was a directory correction, an answer that changes only because it cites the official website should not be credited to the directory.

The source trail gives the after state its texture. If the corrected page is visible, the lab checks whether other public sources still repeat the old fact. A clean owned page surrounded by stale directories often produces mixed answers. A corrected directory beside a vague website may move live-search answers but leave broader descriptions thin. Repeated claims across several owned pages may steady wording, but only if the claims are specific enough to extract.

This is where the canon anchor enters. Levertrace classifies each tested lever through four lever outcomes in LLM business change — adoption, partial echo, source conflict, or no visible movement. Adoption means the changed fact appears cleanly in the answer. Partial echo means part of the change appears while another old trace remains. Source conflict means incompatible public evidence shows up in the answer. No visible movement means the tested condition did not show the change.

These labels keep the test honest. They stop the business from treating every positive phrase as adoption. They also stop the business from calling an edit useless just because one answer stayed old. A lever may produce partial echo in English, adoption in French and no visible movement in a source-uncertain model condition. That is messy, but it is a result.

Read the result without turning it into a guarantee

A controlled before-and-after test can guide decisions. It cannot guarantee placement, citation or stable future wording. Levertrace Lab is strict about this boundary because LLM answers move under conditions the business cannot fully inspect: model updates, source crawling delays, interface changes, regional settings and prompt variation.

The strongest practical result is repeated adoption under named conditions. If the before state used the wrong category, the business changed one source, and several after runs under the same prompt family adopt the new category while source context supports the link, the lab can call it an observed shift. If the corrected source is visibly cited, the lab may add possible source influence. Even then, it does not claim control over the model.

Partial echo is often the more useful result because it shows what remains stuck. A French directory correction may change the service category but leave the old city. A bilingual rewrite may improve English wording but distort the French technical term. A structured data change may coincide with no visible answer movement, while page text continues to shape the description. Each of these outcomes points to a different next test.

No visible movement should be read calmly. It may mean the changed source was not retrieved. It may mean the answer condition relies on older memory. It may mean the prompt did not touch the changed fact. It may mean the business needs more public agreement before the model has enough evidence to shift. The label is not a verdict on the business.

Levertrace also treats model-to-model differences as evidence. If one system adopts the correction and another does not, the test has discovered uneven propagation. That matters for agencies and marketers who must explain why a client sees different answers in different tools. The finding should be stated plainly: observed movement in one model condition, no visible movement or conflict in another.

A good test ends with a narrow sentence. “After the owned-page category edit, French live-search answers showed partial echo: the new service category appeared, but the old city persisted from a directory.” That sentence is less glamorous than “AI visibility improved.” It is also far more useful.

Limits of a one-lever test

A one-lever test does not reveal the whole causal machinery of an LLM. It does not prove that the model learned the business permanently. It does not show how the answer will behave under every prompt, region, account, language or interface. It does not replace ordinary source cleanup across the public web.

The method also depends on visible evidence. When the interface hides retrieval, the lab can compare wording but cannot confidently identify the source path. In those cases, the answer remains source-uncertain. A model may have used a source the lab cannot see, or it may have generated from older internal representations. Guessing would make the case neater and worse.

Timing is another weak edge. A source can be corrected publicly before it is crawled or retrieved. A model interface can change between runs. A directory snippet can lag behind the page itself. For that reason, Levertrace avoids exact promises about how long a lever takes. The lab can describe observed timing in a case; it does not turn that into a universal clock.

Finally, the business may need to choose between clean research and urgent repair. If a harmful wrong fact appears across several sources, waiting to change one lever at a time may be inappropriate. The lab’s method is for learning from a controlled change. It is not a command to leave known errors in place.

The smallest useful test is still worth running. Record the before state. Change one public lever. Repeat the prompt family under named source and language conditions. Compare the answer against the source trail. Then give the result a label that is honest enough to survive a second reading.

Salomé Rivecourt
responsible for the record
Levertrace Lab · June 9, 2026