Tool execution log · T-01–T-04
T-01 | retry=true | stable_key=false | writes=2 T-02 | retry=true | stable_key=true | writes=1 T-03 | retry=false | stable_key=false | writes=1 T-04 | retry=true | stable_key=false | writes=2
02 / Logic discovery
Help AI investigate unfamiliar problems: follow a clue, propose an explanation, seek counterexamples and shape a rule worth reusing.
Start with this exampleYour AI can investigate what comes next.
EXAMPLE 01 / AGENT RELIABILITY
An agent retries a tool call. Some requests produce two writes; others produce one. Follow the investigation from an initial hypothesis to a more precise safeguard.
This walkthrough illustrates a research workflow: discover, challenge, refine and review.
Compare four request records
| Request | Retry | Same key | Writes |
|---|---|---|---|
| T-01 | Yes | No | 2 |
| T-02 | Yes | Yes | 1 |
| T-03 | No | No | 1 |
| T-04 | Yes | No | 2 |
A stable request key tells the tool that two attempts belong to the same request. It helps prevent the same action from being performed twice.
A retry without the same request key is a candidate signal of duplicate risk.
retry ∧ missing_stable_key → review_duplicate_riskT-05No stable key · 2 writes✓T-06Stable key · 1 write✓01 / Notice the anomaly
T-01 and T-04 record duplicate writes. Their common context provides a starting point.
02 / Ask a new question
Compare the retry traces, the request identity and the recorded side effects.
03 / Propose a hypothesis
Test this broad explanation against all four traces, including the cases that worked.
04 / Find a counterexample
T-02 retries with a stable request key and records one write. The broad explanation does not fit.
05 / Refine the logic
The duplicate traces combine a retry with no stable key. That combination becomes a candidate safeguard to test.
06 / Prepare it for reuse
Two further example traces support the refined distinction. Package the candidate, its examples and scope for review.
T-01 and T-04 record duplicate writes. Their common context provides a starting point.
Compare the retry traces, the request identity and the recorded side effects.
Test this broad explanation against all four traces, including the cases that worked.
T-02 retries with a stable request key and records one write. The broad explanation does not fit.
The duplicate traces combine a retry with no stable key. That combination becomes a candidate safeguard to test.
Two further example traces support the refined distinction. Package the candidate, its examples and scope for review.
The investigation produces a clearer question, an evidence-backed explanation and a condition that can be tested again.
T-01 | retry=true | stable_key=false | writes=2 T-02 | retry=true | stable_key=true | writes=1 T-03 | retry=false | stable_key=false | writes=1 T-04 | retry=true | stable_key=false | writes=2
Scope: this example tool writes a record; retries must carry the same idempotency key. The reusable output is a check for duplicate risk.
EXAMPLE 02 / FINDING THE HIDDEN QUESTION
Overall conversion rises from 55% to 78.6%. Looking at each channel reveals a different pattern.
New question: is the improvement broad-based, or explained by who arrived?
| Channel | Before | After |
|---|---|---|
| Organic | 90%90 / 100 | 85%170 / 200 |
| Paid | 20%20 / 100 | 15%3 / 20 |
Each cell shows conversions / visits. Both channel rates fall by 5 percentage points.
Overall conversion
The higher-converting channel accounts for a larger share of visits. Both channel conversion rates fell by 5 percentage points.
A change in the mix masks a decline within both channels.
compare subgroup rates + compare subgroup weights → investigate a mix effect
EXAMPLE 03 / CROSS-SOURCE DISCOVERY
A release note announces a feature. Test results show it working in staging. The production configuration supplies the missing distinction.
Discover a consistency check across documents, environments and tool outputs.
The announcement is published.
The test passed in staging.
Production access is still off.
production_available ← enabled_in(production)
Evidence → Reasoning → Reusable knowledge
Use the available evidence and relevant knowledge to ask questions that were not written into a checklist.
Follow the premises a hypothesis needs. Seek missing facts and evidence that could distinguish competing explanations.
Compare with counterexamples, make the conditions explicit and carry the evidence into review and reuse.
CITPROOF
Bring us a workflow where answers, actions or discoveries need a stronger basis. Explore how evidence chains and logic verification can fit your product.
Talk to the team