Documentation Index

Fetch the complete documentation index at: https://docs.sestek.com/llms.txt

Use this file to discover all available pages before exploring further.

AI Agents Optimization

Prev Next

Optimization reads what a test run found and proposes changes to the agents that produced it. You review the proposals one by one, and nothing reaches an agent's configuration until you apply it.

There are two ways to improve an agent with AI and this is one of them. Optimization starts from a test run, so it argues from evidence about conversations that actually happened. Agent Improver starts from a sentence you write, needs no run first, and opens from the agent's own card. Both propose rather than apply.

Building and running the test is covered in Creating a Test and Reading Test Results. What the module measures overall is on AI Testing.


Starting an optimization

Open a completed run's detail page and select Analyze & Suggest Improvements, beside the run title. It offers two modes.

Mode What it does Availability
Human Approval Proposes changes for you to review and approve before they are applied Available
Self-Optimizing Applies optimizations without review Coming soon

The Analyze & Suggest Improvements menu on a run detail

Self-Optimizing cannot be selected yet, so every optimization goes through Human Approval.

The button appears on runs of a simulated test only. A historical run has no scenario or persona behind it to change, so its detail page does not offer optimization at all.

The evaluation is read in full, not only its failures, so a proposal can protect something the run showed was working as well as fix something it showed was not.


The Human Approval panel

Human Approval opens a panel at the right of the screen. It takes a few moments to read the run, then answers with a summary line and one or more suggestion cards. It proposes something even when the run scored full marks, so a suggestion is not evidence that anything failed.

The Human Approval panel with its suggestion cards

Each card carries a checkbox, ticked by default, a title, a chip naming the agent it applies to, a chip naming what kind of change it is, such as Agent Instructions, and a description of the edit in prose. Apply Selected at the foot of the panel counts what is currently ticked.

Read the description rather than the title. The title says what area is being changed; the description is the actual edit, and it is the only thing you are approving.

Refining a suggestion

The input at the foot of the panel takes an instruction about the proposals themselves, for example asking for them to be shorter or more specific.

The panel after a refinement instruction

A refinement produces a fresh proposal rather than editing the old one. The summary line updates, the previous card is greyed out and unticked, and the new card appears below it. Both stay on screen, so you can compare them and tick whichever you prefer. Nothing has been written to any agent at this point, so an unwanted proposal can simply be left unticked.

The input row at the foot of the panel also takes a file, and the instruction can be dictated instead of typed. Both work the same way in all three panels and are covered on Agent Builder and Improver.


Applying the changes

A card that changes an agent's instructions replaces that field, and the wording it replaces is archived for you on the way past. The previous version stays readable in the agent's Instructions History, labelled as auto-archived, so an applied change can be read back and compared. Archiving is a tenant setting, Automatically archive instructions before each change is applied, on the AI Testing settings page. It is on unless somebody turned it off, so confirm it before you rely on it.

Apply Selected then commits the ticked cards to the agents, and the panel confirms when the configurations have been updated.

Re-run Test is offered beside the confirmation. It runs the same test again, with the same scenarios and personas, so the effect of what you just applied is measured against the same baseline. Expect the figure to move a little on its own between runs, as covered in Reading Test Results.