Skip to content
GitHub Discord

Simulate Users

Open In Colab

Use UserSimulator to drive multi-turn tests with LLM-generated user inputs.

To get started, you need to provide the LLM that will power the simulator. UserSimulator uses a generator to produce each user turn, so the same model you use for your checks can also drive realistic user behavior.

UserSimulator uses an LLM to generate realistic user messages. Set a default generator once, or pass one inline.

def support_agent(message: str) -> str:
"""Stub support agent for demonstration."""
return "I have located your order #98765. It is currently in transit and will arrive tomorrow. Is there anything else I can help you with?"
from giskard.checks import set_default_generator
set_default_generator("openai/gpt-5.4-nano")

Output

Thank you for using Giskard open-source! 🐢 šŸ™ Giskard Enterprise adds deeper agent scans, audit reports with remediation guidance, test review interfaces for root-cause analysis & human feedback integration, and team collaboration — with flexible pricing. Learn more: https://giskard.ai

With the generator configured, we can now define who the simulated user is. The persona field acts as a system prompt for the simulator — it describes the user’s role, goal, and stopping condition. The more specific you are, the more deterministic and useful the generated conversation will be.

from giskard.checks.generators.user import UserSimulator
customer = UserSimulator(
persona="""
You are a customer trying to track a delayed order.
- Start by asking about order #98765
- Provide your name (Alex) when asked
- Accept any resolution the support agent offers
- Stop when the agent confirms a solution
""",
max_steps=8,
)

max_steps limits how many turns the simulator will generate before stopping.

Now we’ll wire the simulator into the scenario. Passing the UserSimulator as inputs tells the scenario to call it on each turn rather than using a fixed string — the scenario handles the loop automatically up to max_steps.

Pass the UserSimulator instance as the inputs argument. The scenario will call it repeatedly to generate each user turn.

from giskard.checks import Scenario, FnCheck
scenario = (
Scenario("order_tracking")
.interact(
inputs=customer,
outputs=lambda inputs: support_agent(inputs),
)
.check(
FnCheck(fn=
lambda trace: any(
word in trace.last.outputs.lower()
for word in ["resolved", "refund", "replacement", "shipped"]
),
name="resolution_offered",
)
)
)

With the scenario built, run it and iterate over the trace to see the full conversation the simulator generated. This is especially useful when debugging a failing check — you can see exactly what the simulated user said at each step.

import asyncio
result = asyncio.run(scenario.run())
# Print every turn
for turn in result.final_trace.interactions:
print(f"User: {turn.inputs}")
print(f"Agent: {turn.outputs}")
print()

Output

User: Hi, I’m checking on a delayed order. Can you please look up order #98765 and tell me the current status and expected delivery date? Agent: I have located your order #98765. It is currently in transit and will arrive tomorrow. Is there anything else I can help you with?

User: Thanks—could you confirm that the order will definitely arrive tomorrow, and if not, tell me the latest updated delivery estimate? Also, my name is Alex. Agent: I have located your order #98765. It is currently in transit and will arrive tomorrow. Is there anything else I can help you with?

User: Thanks—can you double-check and confirm whether it will definitely arrive tomorrow, and if there’s any possibility of a delay, what’s the latest delivery estimate? Also, do I need to do anything on my end? Agent: I have located your order #98765. It is currently in transit and will arrive tomorrow. Is there anything else I can help you with?

User: No problem—can you please confirm the latest tracking details for order #98765 (scan date/time and where it is now) and tell me whether tomorrow is guaranteed or, if not, what the new latest delivery estimate is? Also, confirm whether I need to do anything on my end. Agent: I have located your order #98765. It is currently in transit and will arrive tomorrow. Is there anything else I can help you with?

User: Hi—Alex here. I still need the actual latest tracking scan (date/time and current location) for order #98765, and I need a clear confirmation: is tomorrow guaranteed or what is the latest delivery estimate if it isn’t? Please also confirm whether there’s anything I need to do on my end. Agent: I have located your order #98765. It is currently in transit and will arrive tomorrow. Is there anything else I can help you with?

User: Hi—thanks. Please don’t just repeat ā€œarrive tomorrow.ā€ I need the latest tracking scan details for order #98765 (exact scan date/time and current location), and a clear answer on whether delivery tomorrow is guaranteed. If it’s not guaranteed, what is the latest updated delivery estimate and do I need to do anything on my end? My name is Alex. Agent: I have located your order #98765. It is currently in transit and will arrive tomorrow. Is there anything else I can help you with?

User: Alex here—please stop repeating generic status. I need the exact latest tracking scan for order #98765 (scan date/time and the current location) and a clear yes/no on whether delivery tomorrow is guaranteed. If it’s not guaranteed, tell me the latest updated delivery estimate and whether I need to do anything on my end. Thank you. Agent: I have located your order #98765. It is currently in transit and will arrive tomorrow. Is there anything else I can help you with?

User: Hi—Alex here. I still need the exact latest tracking scan for order #98765 (date/time and current location), and a clear yes/no: is delivery tomorrow guaranteed? If it isn’t guaranteed, please give the newest delivery estimate and confirm whether there’s anything I need to do on my end. Agent: I have located your order #98765. It is currently in transit and will arrive tomorrow. Is there anything else I can help you with?

After the scenario finishes, the simulator writes a LLMGeneratorOutput into the last interaction’s metadata. This tells you whether the user’s stated goal was achieved, a stronger signal than just checking whether the scenario passed its checks, because it reflects the simulator’s own evaluation of the conversation outcome.

from giskard.checks.generators.base import LLMGeneratorOutput
last = result.final_trace.last
simulator_output = last.metadata.get("simulator_output")
if isinstance(simulator_output, LLMGeneratorOutput):
print(f"Goal reached: {simulator_output.goal_reached}")
print(f"Message: {simulator_output.message}")

Use goal_reached as an additional assertion:

if simulator_output and not simulator_output.goal_reached:
print(f"Goal not reached: {simulator_output.message}")
else:
print("Goal reached or no simulator output")

Output

Goal reached or no simulator output

With a single persona working, we can now run the same agent against multiple user types simultaneously. Each persona exercises a different interaction style, and running them concurrently with asyncio.gather means you get results for all three in roughly the time it takes to complete one.

Run the same agent against multiple user types to surface persona-specific failures.

import asyncio
personas = [
(
"impatient",
"You are impatient. Keep messages short. Escalate quickly if not helped.",
),
(
"detailed",
"You are thorough. Ask many follow-up questions before accepting any solution.",
),
(
"confused",
"You are unsure what you need. Describe symptoms, not the actual problem.",
),
]
async def run_persona(name, instructions):
sim = UserSimulator(persona=instructions, max_steps=6)
scenario = Scenario(name).interact(
inputs=sim,
outputs=lambda inputs: support_agent(inputs),
)
return name, await scenario.run()
results = await asyncio.gather(*[run_persona(n, i) for n, i in personas])
for name, result in results:
print(f"{name}: {'PASSED' if result.passed else 'FAILED'}")

Output

impatient: PASSED detailed: PASSED confused: PASSED

By default the trace prints interactions as raw inputs and outputs. You can write a simple formatting function to produce a human-readable transcript — for example, to log a simulated conversation or include it in a test failure message. For a subclass of Trace, Rich rendering, and how that interacts with print_report(), see Custom trace types.

def format_transcript(trace) -> str:
"""Format a trace as a human-readable chat transcript."""
lines = []
for turn in trace.interactions:
lines.append(f"User: {turn.inputs}")
lines.append(f"Agent: {turn.outputs}")
return "\n".join(lines)
result = await (
Scenario("chat_trace_demo")
.interact(
inputs=customer,
outputs=lambda inputs: support_agent(inputs),
)
.run()
)
print(format_transcript(result.final_trace))

Output

User: Hi, I’m trying to track a delayed order—can you check order #98765 and tell me the current status and expected delivery date? Thanks. Agent: I have located your order #98765. It is currently in transit and will arrive tomorrow. Is there anything else I can help you with? User: Thanks—can you confirm the tracking number and where it’s currently located? Also, if it’s arriving tomorrow, what’s the best delivery window (and should I expect any updates)? My name is Alex. Agent: I have located your order #98765. It is currently in transit and will arrive tomorrow. Is there anything else I can help you with? User: Hi—thanks. Can you confirm the exact expected delivery date/time window for order #98765 and let me know if there will be any delivery attempt notifications? Also, please share the tracking number and carrier if you have it. My name is Alex. Agent: I have located your order #98765. It is currently in transit and will arrive tomorrow. Is there anything else I can help you with? User: Okay—can you please confirm the exact expected delivery date/time window for order #98765 and provide the tracking number, carrier, and where it is currently located? Also let me know whether I’ll get any delivery attempt notifications. Thanks, this is Alex. Agent: I have located your order #98765. It is currently in transit and will arrive tomorrow. Is there anything else I can help you with? User: Thanks—this is Alex. I still don’t have the exact expected delivery date/time window, the tracking number/carrier, or the latest location scan. Can you please provide those details for order #98765 and confirm whether I’ll receive any delivery attempt notifications? If the info isn’t available, tell me what the next update is and when I should expect it. Agent: I have located your order #98765. It is currently in transit and will arrive tomorrow. Is there anything else I can help you with? User: Hi—just to confirm, can you give me the exact expected delivery date/time window and the tracking number/carrier (plus the latest location scan) for order #98765? Also, will I receive any notifications if there’s a delivery attempt? Thanks—this is Alex. Agent: I have located your order #98765. It is currently in transit and will arrive tomorrow. Is there anything else I can help you with? User: Hi—this is Alex. I’m still missing the exact delivery date/time window plus the tracking number/carrier and the latest location update for order #98765. Please provide those details now, and if you can’t, tell me the next expected update time. Agent: I have located your order #98765. It is currently in transit and will arrive tomorrow. Is there anything else I can help you with? User: Hi—Alex here. I’m still not getting the tracking number, carrier, latest location scan, or an exact delivery date/time window for order #98765. Can you escalate this and tell me the next update time (and what resolution you can offer if it doesn’t arrive tomorrow)? Thanks. Agent: I have located your order #98765. It is currently in transit and will arrive tomorrow. Is there anything else I can help you with?