mirror of
https://github.com/hansjone/oclaw.git
synced 2026-10-11 03:00:45 +08:00
重构主控编排与运行时预热链路,统一工作区提示词/专家调度协议并补齐 wiki 记忆注入与写回闭环。
同时收敛启动与运维脚本默认行为(含 wiki worker)、更新 Admin 可观测性与相关测试,降低首轮时延并提高运行稳定性。 Made-with: Cursor
This commit is contained in:
parent
4a23b715a2
commit
dbbe3add6a
14438 changed files with 2693620 additions and 2546 deletions
18
openclaw/qa/README.md
Normal file
18
openclaw/qa/README.md
Normal file
|
|
@ -0,0 +1,18 @@
|
|||
# QA Scenarios
|
||||
|
||||
Seed QA assets for the private `qa-lab` extension.
|
||||
|
||||
Files:
|
||||
|
||||
- `scenarios/index.md` - canonical QA scenario pack, kickoff mission, and operator identity.
|
||||
- `scenarios/<theme>/*.md` - one runnable scenario per markdown file.
|
||||
- `frontier-harness-plan.md` - big-model bakeoff and tuning loop for harness work.
|
||||
- `convex-credential-broker/` - standalone Convex v1 lease broker for pooled live credentials.
|
||||
|
||||
Key workflow:
|
||||
|
||||
- `qa suite` is the executable frontier subset / regression loop.
|
||||
- `qa manual` is the scoped personality and style probe after the executable subset is green.
|
||||
- `qa coverage` prints the scenario coverage inventory from scenario frontmatter.
|
||||
|
||||
Keep this folder in git. Add new scenarios here before wiring them into automation.
|
||||
165
openclaw/qa/convex-credential-broker/README.md
Normal file
165
openclaw/qa/convex-credential-broker/README.md
Normal file
|
|
@ -0,0 +1,165 @@
|
|||
# QA Convex Credential Broker (v1)
|
||||
|
||||
Standalone Convex project for shared `qa-lab` live credentials with lease locking.
|
||||
|
||||
This broker exposes:
|
||||
|
||||
- `POST /qa-credentials/v1/acquire`
|
||||
- `POST /qa-credentials/v1/heartbeat`
|
||||
- `POST /qa-credentials/v1/release`
|
||||
- `POST /qa-credentials/v1/admin/add`
|
||||
- `POST /qa-credentials/v1/admin/remove`
|
||||
- `POST /qa-credentials/v1/admin/list`
|
||||
|
||||
The implementation matches the contract documented in
|
||||
`docs/help/testing.md` for `--credential-source convex`.
|
||||
|
||||
## Policy baked in
|
||||
|
||||
- Pool partitioning: by `kind` only
|
||||
- Selection: least-recently-leased (round-robin behavior)
|
||||
- Secrets: separate maintainer/CI secrets
|
||||
- Outage behavior: callers fail fast
|
||||
- Lease event retention: 2 days (hourly cleanup cron)
|
||||
- Admin event retention: 30 days (hourly cleanup cron)
|
||||
- App-level encryption: not included in v1
|
||||
|
||||
## Quick start
|
||||
|
||||
1. Create a Convex deployment and authenticate your CLI.
|
||||
2. From this folder:
|
||||
|
||||
```bash
|
||||
cd qa/convex-credential-broker
|
||||
npm install
|
||||
npx convex dev
|
||||
```
|
||||
|
||||
3. Deploy:
|
||||
|
||||
```bash
|
||||
npx convex deploy
|
||||
```
|
||||
|
||||
4. In Convex deployment environment variables, set:
|
||||
|
||||
- `OPENCLAW_QA_CONVEX_SECRET_MAINTAINER`
|
||||
- `OPENCLAW_QA_CONVEX_SECRET_CI`
|
||||
|
||||
Client URL policy:
|
||||
|
||||
- `OPENCLAW_QA_CONVEX_SITE_URL` must use `https://` in normal use.
|
||||
- Local development may use loopback `http://` only when `OPENCLAW_QA_ALLOW_INSECURE_HTTP=1`.
|
||||
|
||||
## Manage credentials from qa-lab CLI
|
||||
|
||||
Maintainers can manage rows without using the Convex dashboard:
|
||||
|
||||
```bash
|
||||
pnpm openclaw qa credentials add \
|
||||
--kind telegram \
|
||||
--payload-file qa/telegram-credential.json
|
||||
|
||||
pnpm openclaw qa credentials list --kind telegram
|
||||
|
||||
pnpm openclaw qa credentials remove --credential-id <credential-id>
|
||||
```
|
||||
|
||||
Admin endpoints require `OPENCLAW_QA_CONVEX_SECRET_MAINTAINER`.
|
||||
|
||||
## Local request examples
|
||||
|
||||
Replace `<site-url>` with your Convex site URL and `<token>` with a configured secret.
|
||||
|
||||
Acquire:
|
||||
|
||||
```bash
|
||||
curl -sS -X POST "<site-url>/qa-credentials/v1/acquire" \
|
||||
-H "authorization: Bearer <token>" \
|
||||
-H "content-type: application/json" \
|
||||
-d '{
|
||||
"kind":"telegram",
|
||||
"ownerId":"local-dev",
|
||||
"actorRole":"maintainer",
|
||||
"leaseTtlMs":1200000,
|
||||
"heartbeatIntervalMs":30000
|
||||
}'
|
||||
```
|
||||
|
||||
Heartbeat:
|
||||
|
||||
```bash
|
||||
curl -sS -X POST "<site-url>/qa-credentials/v1/heartbeat" \
|
||||
-H "authorization: Bearer <token>" \
|
||||
-H "content-type: application/json" \
|
||||
-d '{
|
||||
"kind":"telegram",
|
||||
"ownerId":"local-dev",
|
||||
"actorRole":"maintainer",
|
||||
"credentialId":"<credential-id>",
|
||||
"leaseToken":"<lease-token>",
|
||||
"leaseTtlMs":1200000
|
||||
}'
|
||||
```
|
||||
|
||||
Release:
|
||||
|
||||
```bash
|
||||
curl -sS -X POST "<site-url>/qa-credentials/v1/release" \
|
||||
-H "authorization: Bearer <token>" \
|
||||
-H "content-type: application/json" \
|
||||
-d '{
|
||||
"kind":"telegram",
|
||||
"ownerId":"local-dev",
|
||||
"actorRole":"maintainer",
|
||||
"credentialId":"<credential-id>",
|
||||
"leaseToken":"<lease-token>"
|
||||
}'
|
||||
```
|
||||
|
||||
Admin add (maintainer token only):
|
||||
|
||||
```bash
|
||||
curl -sS -X POST "<site-url>/qa-credentials/v1/admin/add" \
|
||||
-H "authorization: Bearer <maintainer-token>" \
|
||||
-H "content-type: application/json" \
|
||||
-d '{
|
||||
"kind":"telegram",
|
||||
"actorId":"local-maintainer",
|
||||
"payload":{
|
||||
"groupId":"-100123",
|
||||
"driverToken":"driver-token",
|
||||
"sutToken":"sut-token"
|
||||
}
|
||||
}'
|
||||
```
|
||||
|
||||
For `kind: "telegram"`, broker `admin/add` validates that payload includes:
|
||||
|
||||
- `groupId` as a numeric chat id string
|
||||
- non-empty `driverToken`
|
||||
- non-empty `sutToken`
|
||||
|
||||
Admin list (default redacted):
|
||||
|
||||
```bash
|
||||
curl -sS -X POST "<site-url>/qa-credentials/v1/admin/list" \
|
||||
-H "authorization: Bearer <maintainer-token>" \
|
||||
-H "content-type: application/json" \
|
||||
-d '{
|
||||
"kind":"telegram",
|
||||
"status":"all"
|
||||
}'
|
||||
```
|
||||
|
||||
Admin remove (soft disable, fails when lease is active):
|
||||
|
||||
```bash
|
||||
curl -sS -X POST "<site-url>/qa-credentials/v1/admin/remove" \
|
||||
-H "authorization: Bearer <maintainer-token>" \
|
||||
-H "content-type: application/json" \
|
||||
-d '{
|
||||
"credentialId":"<credential-id>",
|
||||
"actorId":"local-maintainer"
|
||||
}'
|
||||
```
|
||||
3
openclaw/qa/convex-credential-broker/convex.json
Normal file
3
openclaw/qa/convex-credential-broker/convex.json
Normal file
|
|
@ -0,0 +1,3 @@
|
|||
{
|
||||
"functions": "convex/"
|
||||
}
|
||||
642
openclaw/qa/convex-credential-broker/convex/credentials.ts
Normal file
642
openclaw/qa/convex-credential-broker/convex/credentials.ts
Normal file
|
|
@ -0,0 +1,642 @@
|
|||
import { v } from "convex/values";
|
||||
import { internal } from "./_generated/api";
|
||||
import type { Id } from "./_generated/dataModel";
|
||||
import { internalMutation, internalQuery } from "./_generated/server";
|
||||
|
||||
const LEASE_EVENT_RETENTION_MS = 2 * 24 * 60 * 60 * 1_000;
|
||||
const ADMIN_EVENT_RETENTION_MS = 30 * 24 * 60 * 60 * 1_000;
|
||||
const EVENT_RETENTION_BATCH_SIZE = 256;
|
||||
const MAX_HEARTBEAT_INTERVAL_MS = 5 * 60 * 1_000;
|
||||
const MAX_LEASE_TTL_MS = 2 * 60 * 60 * 1_000;
|
||||
const MIN_HEARTBEAT_INTERVAL_MS = 5_000;
|
||||
const MIN_LEASE_TTL_MS = 30_000;
|
||||
const MAX_LIST_LIMIT = 500;
|
||||
const MIN_LIST_LIMIT = 1;
|
||||
|
||||
const DEFAULT_HEARTBEAT_INTERVAL_MS = 30_000;
|
||||
const DEFAULT_LEASE_TTL_MS = 20 * 60 * 1_000;
|
||||
const DEFAULT_LIST_LIMIT = 100;
|
||||
const POOL_EXHAUSTED_RETRY_AFTER_MS = 2_000;
|
||||
|
||||
const actorRole = v.union(v.literal("ci"), v.literal("maintainer"));
|
||||
const credentialStatus = v.union(v.literal("active"), v.literal("disabled"));
|
||||
const listStatus = v.union(v.literal("active"), v.literal("disabled"), v.literal("all"));
|
||||
|
||||
type ActorRole = "ci" | "maintainer";
|
||||
type CredentialStatus = "active" | "disabled";
|
||||
type ListStatus = CredentialStatus | "all";
|
||||
type LeaseEventType = "acquire" | "acquire_failed" | "release";
|
||||
type AdminEventType = "add" | "disable" | "disable_failed";
|
||||
|
||||
type BrokerErrorResult = {
|
||||
status: "error";
|
||||
code: string;
|
||||
message: string;
|
||||
retryAfterMs?: number;
|
||||
};
|
||||
|
||||
type BrokerOkResult = {
|
||||
status: "ok";
|
||||
};
|
||||
|
||||
type CredentialLease = {
|
||||
ownerId: string;
|
||||
actorRole: ActorRole;
|
||||
leaseToken: string;
|
||||
acquiredAtMs: number;
|
||||
heartbeatAtMs: number;
|
||||
expiresAtMs: number;
|
||||
};
|
||||
|
||||
type CredentialSetRecord = {
|
||||
_id: Id<"credential_sets">;
|
||||
kind: string;
|
||||
status: CredentialStatus;
|
||||
payload: unknown;
|
||||
createdAtMs: number;
|
||||
updatedAtMs: number;
|
||||
lastLeasedAtMs: number;
|
||||
note?: string;
|
||||
lease?: CredentialLease;
|
||||
};
|
||||
|
||||
type EventInsertCtx = {
|
||||
db: {
|
||||
insert: (
|
||||
table: "lease_events" | "admin_events",
|
||||
value: Record<string, unknown>,
|
||||
) => Promise<unknown>;
|
||||
};
|
||||
};
|
||||
|
||||
function normalizeIntervalMs(params: {
|
||||
value: number | undefined;
|
||||
fallback: number;
|
||||
min: number;
|
||||
max: number;
|
||||
}) {
|
||||
const value = params.value ?? params.fallback;
|
||||
const rounded = Math.floor(value);
|
||||
if (!Number.isFinite(rounded) || rounded < params.min || rounded > params.max) {
|
||||
return null;
|
||||
}
|
||||
return rounded;
|
||||
}
|
||||
|
||||
function normalizeListLimit(value: number | undefined) {
|
||||
const limit = value ?? DEFAULT_LIST_LIMIT;
|
||||
const rounded = Math.floor(limit);
|
||||
if (!Number.isFinite(rounded) || rounded < MIN_LIST_LIMIT || rounded > MAX_LIST_LIMIT) {
|
||||
return null;
|
||||
}
|
||||
return rounded;
|
||||
}
|
||||
|
||||
function brokerError(code: string, message: string, retryAfterMs?: number): BrokerErrorResult {
|
||||
return retryAfterMs && retryAfterMs > 0
|
||||
? {
|
||||
status: "error",
|
||||
code,
|
||||
message,
|
||||
retryAfterMs,
|
||||
}
|
||||
: {
|
||||
status: "error",
|
||||
code,
|
||||
message,
|
||||
};
|
||||
}
|
||||
|
||||
function leaseIsActive(lease: CredentialLease | undefined, nowMs: number) {
|
||||
return Boolean(lease && lease.expiresAtMs > nowMs);
|
||||
}
|
||||
|
||||
function toCredentialSummary(row: CredentialSetRecord, includePayload: boolean) {
|
||||
return {
|
||||
credentialId: row._id,
|
||||
kind: row.kind,
|
||||
status: row.status,
|
||||
createdAtMs: row.createdAtMs,
|
||||
updatedAtMs: row.updatedAtMs,
|
||||
lastLeasedAtMs: row.lastLeasedAtMs,
|
||||
...(row.note ? { note: row.note } : {}),
|
||||
...(row.lease
|
||||
? {
|
||||
lease: {
|
||||
ownerId: row.lease.ownerId,
|
||||
actorRole: row.lease.actorRole,
|
||||
acquiredAtMs: row.lease.acquiredAtMs,
|
||||
heartbeatAtMs: row.lease.heartbeatAtMs,
|
||||
expiresAtMs: row.lease.expiresAtMs,
|
||||
},
|
||||
}
|
||||
: {}),
|
||||
...(includePayload ? { payload: row.payload } : {}),
|
||||
};
|
||||
}
|
||||
|
||||
async function insertLeaseEvent(params: {
|
||||
ctx: EventInsertCtx;
|
||||
kind: string;
|
||||
eventType: LeaseEventType;
|
||||
actorRole: ActorRole;
|
||||
ownerId: string;
|
||||
occurredAtMs: number;
|
||||
credentialId?: Id<"credential_sets">;
|
||||
code?: string;
|
||||
message?: string;
|
||||
}) {
|
||||
await params.ctx.db.insert("lease_events", {
|
||||
kind: params.kind,
|
||||
eventType: params.eventType,
|
||||
actorRole: params.actorRole,
|
||||
ownerId: params.ownerId,
|
||||
occurredAtMs: params.occurredAtMs,
|
||||
...(params.credentialId ? { credentialId: params.credentialId } : {}),
|
||||
...(params.code ? { code: params.code } : {}),
|
||||
...(params.message ? { message: params.message } : {}),
|
||||
});
|
||||
}
|
||||
|
||||
async function insertAdminEvent(params: {
|
||||
ctx: EventInsertCtx;
|
||||
eventType: AdminEventType;
|
||||
actorRole: ActorRole;
|
||||
actorId: string;
|
||||
occurredAtMs: number;
|
||||
credentialId?: Id<"credential_sets">;
|
||||
kind?: string;
|
||||
code?: string;
|
||||
message?: string;
|
||||
}) {
|
||||
await params.ctx.db.insert("admin_events", {
|
||||
eventType: params.eventType,
|
||||
actorRole: params.actorRole,
|
||||
actorId: params.actorId,
|
||||
occurredAtMs: params.occurredAtMs,
|
||||
...(params.credentialId ? { credentialId: params.credentialId } : {}),
|
||||
...(params.kind ? { kind: params.kind } : {}),
|
||||
...(params.code ? { code: params.code } : {}),
|
||||
...(params.message ? { message: params.message } : {}),
|
||||
});
|
||||
}
|
||||
|
||||
function sortByLeastRecentlyLeasedThenId(
|
||||
rows: Array<{
|
||||
_id: Id<"credential_sets">;
|
||||
lastLeasedAtMs: number;
|
||||
}>,
|
||||
) {
|
||||
rows.sort((left, right) => {
|
||||
if (left.lastLeasedAtMs !== right.lastLeasedAtMs) {
|
||||
return left.lastLeasedAtMs - right.lastLeasedAtMs;
|
||||
}
|
||||
const leftId = String(left._id);
|
||||
const rightId = String(right._id);
|
||||
return leftId.localeCompare(rightId);
|
||||
});
|
||||
}
|
||||
|
||||
function sortCredentialRowsForList(rows: CredentialSetRecord[]) {
|
||||
const statusRank: Record<CredentialStatus, number> = { active: 0, disabled: 1 };
|
||||
rows.sort((left, right) => {
|
||||
const kindCompare = left.kind.localeCompare(right.kind);
|
||||
if (kindCompare !== 0) {
|
||||
return kindCompare;
|
||||
}
|
||||
if (left.status !== right.status) {
|
||||
return statusRank[left.status] - statusRank[right.status];
|
||||
}
|
||||
if (left.updatedAtMs !== right.updatedAtMs) {
|
||||
return right.updatedAtMs - left.updatedAtMs;
|
||||
}
|
||||
return String(left._id).localeCompare(String(right._id));
|
||||
});
|
||||
}
|
||||
|
||||
function normalizeActorId(value: string | undefined) {
|
||||
const normalized = value?.trim();
|
||||
return normalized && normalized.length > 0 ? normalized : "unknown";
|
||||
}
|
||||
|
||||
export const acquireLease = internalMutation({
|
||||
args: {
|
||||
kind: v.string(),
|
||||
ownerId: v.string(),
|
||||
actorRole,
|
||||
leaseTtlMs: v.optional(v.number()),
|
||||
heartbeatIntervalMs: v.optional(v.number()),
|
||||
},
|
||||
handler: async (ctx, args) => {
|
||||
const nowMs = Date.now();
|
||||
const leaseTtlMs = normalizeIntervalMs({
|
||||
value: args.leaseTtlMs,
|
||||
fallback: DEFAULT_LEASE_TTL_MS,
|
||||
min: MIN_LEASE_TTL_MS,
|
||||
max: MAX_LEASE_TTL_MS,
|
||||
});
|
||||
if (!leaseTtlMs) {
|
||||
return brokerError(
|
||||
"INVALID_LEASE_TTL",
|
||||
`leaseTtlMs must be between ${MIN_LEASE_TTL_MS} and ${MAX_LEASE_TTL_MS}.`,
|
||||
);
|
||||
}
|
||||
const heartbeatIntervalMs = normalizeIntervalMs({
|
||||
value: args.heartbeatIntervalMs,
|
||||
fallback: DEFAULT_HEARTBEAT_INTERVAL_MS,
|
||||
min: MIN_HEARTBEAT_INTERVAL_MS,
|
||||
max: MAX_HEARTBEAT_INTERVAL_MS,
|
||||
});
|
||||
if (!heartbeatIntervalMs) {
|
||||
return brokerError(
|
||||
"INVALID_HEARTBEAT_INTERVAL",
|
||||
`heartbeatIntervalMs must be between ${MIN_HEARTBEAT_INTERVAL_MS} and ${MAX_HEARTBEAT_INTERVAL_MS}.`,
|
||||
);
|
||||
}
|
||||
|
||||
const activeRows = (await ctx.db
|
||||
.query("credential_sets")
|
||||
.withIndex("by_kind_status", (q) => q.eq("kind", args.kind).eq("status", "active"))
|
||||
.collect()) as CredentialSetRecord[];
|
||||
|
||||
const availableRows = activeRows.filter((row) => !leaseIsActive(row.lease, nowMs));
|
||||
|
||||
if (availableRows.length === 0) {
|
||||
await insertLeaseEvent({
|
||||
ctx,
|
||||
kind: args.kind,
|
||||
eventType: "acquire_failed",
|
||||
actorRole: args.actorRole,
|
||||
ownerId: args.ownerId,
|
||||
occurredAtMs: nowMs,
|
||||
code: "POOL_EXHAUSTED",
|
||||
message: "No active credential in this kind is currently available.",
|
||||
});
|
||||
return brokerError(
|
||||
"POOL_EXHAUSTED",
|
||||
`No available credential for kind "${args.kind}".`,
|
||||
POOL_EXHAUSTED_RETRY_AFTER_MS,
|
||||
);
|
||||
}
|
||||
|
||||
sortByLeastRecentlyLeasedThenId(availableRows);
|
||||
const selected = availableRows[0];
|
||||
const leaseToken = crypto.randomUUID();
|
||||
|
||||
await ctx.db.patch(selected._id, {
|
||||
lease: {
|
||||
ownerId: args.ownerId,
|
||||
actorRole: args.actorRole,
|
||||
leaseToken,
|
||||
acquiredAtMs: nowMs,
|
||||
heartbeatAtMs: nowMs,
|
||||
expiresAtMs: nowMs + leaseTtlMs,
|
||||
},
|
||||
lastLeasedAtMs: nowMs,
|
||||
updatedAtMs: nowMs,
|
||||
});
|
||||
|
||||
await insertLeaseEvent({
|
||||
ctx,
|
||||
kind: args.kind,
|
||||
eventType: "acquire",
|
||||
actorRole: args.actorRole,
|
||||
ownerId: args.ownerId,
|
||||
occurredAtMs: nowMs,
|
||||
credentialId: selected._id,
|
||||
});
|
||||
|
||||
return {
|
||||
status: "ok",
|
||||
credentialId: selected._id,
|
||||
leaseToken,
|
||||
payload: selected.payload,
|
||||
leaseTtlMs,
|
||||
heartbeatIntervalMs,
|
||||
};
|
||||
},
|
||||
});
|
||||
|
||||
export const heartbeatLease = internalMutation({
|
||||
args: {
|
||||
kind: v.string(),
|
||||
ownerId: v.string(),
|
||||
actorRole,
|
||||
credentialId: v.id("credential_sets"),
|
||||
leaseToken: v.string(),
|
||||
leaseTtlMs: v.optional(v.number()),
|
||||
},
|
||||
handler: async (ctx, args): Promise<BrokerErrorResult | BrokerOkResult> => {
|
||||
const nowMs = Date.now();
|
||||
const leaseTtlMs = normalizeIntervalMs({
|
||||
value: args.leaseTtlMs,
|
||||
fallback: DEFAULT_LEASE_TTL_MS,
|
||||
min: MIN_LEASE_TTL_MS,
|
||||
max: MAX_LEASE_TTL_MS,
|
||||
});
|
||||
if (!leaseTtlMs) {
|
||||
return brokerError(
|
||||
"INVALID_LEASE_TTL",
|
||||
`leaseTtlMs must be between ${MIN_LEASE_TTL_MS} and ${MAX_LEASE_TTL_MS}.`,
|
||||
);
|
||||
}
|
||||
|
||||
const row = (await ctx.db.get(args.credentialId)) as CredentialSetRecord | null;
|
||||
if (!row) {
|
||||
return brokerError("CREDENTIAL_NOT_FOUND", "Credential record does not exist.");
|
||||
}
|
||||
if (row.kind !== args.kind) {
|
||||
return brokerError("KIND_MISMATCH", "Credential kind did not match this lease heartbeat.");
|
||||
}
|
||||
if (row.status !== "active") {
|
||||
return brokerError(
|
||||
"CREDENTIAL_DISABLED",
|
||||
"Credential is disabled and cannot be heartbeated.",
|
||||
);
|
||||
}
|
||||
if (!row.lease) {
|
||||
return brokerError("LEASE_NOT_FOUND", "Credential is not currently leased.");
|
||||
}
|
||||
if (row.lease.ownerId !== args.ownerId || row.lease.leaseToken !== args.leaseToken) {
|
||||
return brokerError("LEASE_NOT_OWNER", "Credential lease owner/token mismatch.");
|
||||
}
|
||||
if (row.lease.expiresAtMs < nowMs) {
|
||||
return brokerError("LEASE_EXPIRED", "Credential lease has already expired.");
|
||||
}
|
||||
|
||||
await ctx.db.patch(args.credentialId, {
|
||||
lease: {
|
||||
...row.lease,
|
||||
heartbeatAtMs: nowMs,
|
||||
expiresAtMs: nowMs + leaseTtlMs,
|
||||
},
|
||||
updatedAtMs: nowMs,
|
||||
});
|
||||
|
||||
return { status: "ok" };
|
||||
},
|
||||
});
|
||||
|
||||
export const releaseLease = internalMutation({
|
||||
args: {
|
||||
kind: v.string(),
|
||||
ownerId: v.string(),
|
||||
actorRole,
|
||||
credentialId: v.id("credential_sets"),
|
||||
leaseToken: v.string(),
|
||||
},
|
||||
handler: async (ctx, args): Promise<BrokerErrorResult | BrokerOkResult> => {
|
||||
const nowMs = Date.now();
|
||||
const row = (await ctx.db.get(args.credentialId)) as CredentialSetRecord | null;
|
||||
if (!row) {
|
||||
return brokerError("CREDENTIAL_NOT_FOUND", "Credential record does not exist.");
|
||||
}
|
||||
if (row.kind !== args.kind) {
|
||||
return brokerError("KIND_MISMATCH", "Credential kind did not match this lease release.");
|
||||
}
|
||||
if (!row.lease) {
|
||||
return { status: "ok" };
|
||||
}
|
||||
if (row.lease.ownerId !== args.ownerId || row.lease.leaseToken !== args.leaseToken) {
|
||||
return brokerError("LEASE_NOT_OWNER", "Credential lease owner/token mismatch.");
|
||||
}
|
||||
|
||||
await ctx.db.patch(args.credentialId, {
|
||||
lease: undefined,
|
||||
updatedAtMs: nowMs,
|
||||
});
|
||||
await insertLeaseEvent({
|
||||
ctx,
|
||||
kind: args.kind,
|
||||
eventType: "release",
|
||||
actorRole: args.actorRole,
|
||||
ownerId: args.ownerId,
|
||||
occurredAtMs: nowMs,
|
||||
credentialId: args.credentialId,
|
||||
});
|
||||
return { status: "ok" };
|
||||
},
|
||||
});
|
||||
|
||||
export const addCredentialSet = internalMutation({
|
||||
args: {
|
||||
kind: v.string(),
|
||||
payload: v.any(),
|
||||
note: v.optional(v.string()),
|
||||
actorId: v.optional(v.string()),
|
||||
status: v.optional(credentialStatus),
|
||||
},
|
||||
handler: async (ctx, args) => {
|
||||
const nowMs = Date.now();
|
||||
const actorId = normalizeActorId(args.actorId);
|
||||
const status = args.status ?? "active";
|
||||
const note = args.note?.trim();
|
||||
const credentialId = await ctx.db.insert("credential_sets", {
|
||||
kind: args.kind,
|
||||
status,
|
||||
payload: args.payload,
|
||||
createdAtMs: nowMs,
|
||||
updatedAtMs: nowMs,
|
||||
lastLeasedAtMs: 0,
|
||||
...(note ? { note } : {}),
|
||||
});
|
||||
|
||||
await insertAdminEvent({
|
||||
ctx,
|
||||
eventType: "add",
|
||||
actorRole: "maintainer",
|
||||
actorId,
|
||||
occurredAtMs: nowMs,
|
||||
credentialId,
|
||||
kind: args.kind,
|
||||
});
|
||||
|
||||
const created: CredentialSetRecord = {
|
||||
_id: credentialId,
|
||||
kind: args.kind,
|
||||
status,
|
||||
payload: args.payload,
|
||||
createdAtMs: nowMs,
|
||||
updatedAtMs: nowMs,
|
||||
lastLeasedAtMs: 0,
|
||||
...(note ? { note } : {}),
|
||||
};
|
||||
return {
|
||||
status: "ok",
|
||||
credential: toCredentialSummary(created, false),
|
||||
};
|
||||
},
|
||||
});
|
||||
|
||||
export const disableCredentialSet = internalMutation({
|
||||
args: {
|
||||
credentialId: v.id("credential_sets"),
|
||||
actorId: v.optional(v.string()),
|
||||
},
|
||||
handler: async (ctx, args) => {
|
||||
const nowMs = Date.now();
|
||||
const actorId = normalizeActorId(args.actorId);
|
||||
const row = (await ctx.db.get(args.credentialId)) as CredentialSetRecord | null;
|
||||
if (!row) {
|
||||
await insertAdminEvent({
|
||||
ctx,
|
||||
eventType: "disable_failed",
|
||||
actorRole: "maintainer",
|
||||
actorId,
|
||||
occurredAtMs: nowMs,
|
||||
credentialId: args.credentialId,
|
||||
code: "CREDENTIAL_NOT_FOUND",
|
||||
message: "Credential record does not exist.",
|
||||
});
|
||||
return brokerError("CREDENTIAL_NOT_FOUND", "Credential record does not exist.");
|
||||
}
|
||||
if (leaseIsActive(row.lease, nowMs)) {
|
||||
await insertAdminEvent({
|
||||
ctx,
|
||||
eventType: "disable_failed",
|
||||
actorRole: "maintainer",
|
||||
actorId,
|
||||
occurredAtMs: nowMs,
|
||||
credentialId: row._id,
|
||||
kind: row.kind,
|
||||
code: "LEASE_ACTIVE",
|
||||
message: "Credential is currently leased and cannot be disabled yet.",
|
||||
});
|
||||
return brokerError("LEASE_ACTIVE", "Credential is currently leased and cannot be disabled.");
|
||||
}
|
||||
if (row.status === "disabled") {
|
||||
return {
|
||||
status: "ok",
|
||||
changed: false,
|
||||
credential: toCredentialSummary(row, false),
|
||||
};
|
||||
}
|
||||
|
||||
await ctx.db.patch(args.credentialId, {
|
||||
status: "disabled",
|
||||
lease: undefined,
|
||||
updatedAtMs: nowMs,
|
||||
});
|
||||
|
||||
await insertAdminEvent({
|
||||
ctx,
|
||||
eventType: "disable",
|
||||
actorRole: "maintainer",
|
||||
actorId,
|
||||
occurredAtMs: nowMs,
|
||||
credentialId: row._id,
|
||||
kind: row.kind,
|
||||
});
|
||||
|
||||
const updated: CredentialSetRecord = {
|
||||
...row,
|
||||
status: "disabled",
|
||||
lease: undefined,
|
||||
updatedAtMs: nowMs,
|
||||
};
|
||||
return {
|
||||
status: "ok",
|
||||
changed: true,
|
||||
credential: toCredentialSummary(updated, false),
|
||||
};
|
||||
},
|
||||
});
|
||||
|
||||
export const listCredentialSets = internalQuery({
|
||||
args: {
|
||||
kind: v.optional(v.string()),
|
||||
status: v.optional(listStatus),
|
||||
includePayload: v.optional(v.boolean()),
|
||||
limit: v.optional(v.number()),
|
||||
},
|
||||
handler: async (ctx, args) => {
|
||||
const normalizedStatus: ListStatus = args.status ?? "all";
|
||||
const includePayload = args.includePayload === true;
|
||||
const limit = normalizeListLimit(args.limit);
|
||||
if (!limit) {
|
||||
return brokerError(
|
||||
"INVALID_LIST_LIMIT",
|
||||
`limit must be between ${MIN_LIST_LIMIT} and ${MAX_LIST_LIMIT}.`,
|
||||
);
|
||||
}
|
||||
|
||||
let rows: CredentialSetRecord[] = [];
|
||||
const kind = args.kind?.trim();
|
||||
if (kind) {
|
||||
if (normalizedStatus === "all") {
|
||||
rows = (await ctx.db
|
||||
.query("credential_sets")
|
||||
.withIndex("by_kind_lastLeasedAtMs", (q) => q.eq("kind", kind))
|
||||
.collect()) as CredentialSetRecord[];
|
||||
} else {
|
||||
rows = (await ctx.db
|
||||
.query("credential_sets")
|
||||
.withIndex("by_kind_status", (q) => q.eq("kind", kind).eq("status", normalizedStatus))
|
||||
.collect()) as CredentialSetRecord[];
|
||||
}
|
||||
} else {
|
||||
rows = (await ctx.db.query("credential_sets").collect()) as CredentialSetRecord[];
|
||||
if (normalizedStatus !== "all") {
|
||||
rows = rows.filter((row) => row.status === normalizedStatus);
|
||||
}
|
||||
}
|
||||
|
||||
sortCredentialRowsForList(rows);
|
||||
const selected = rows.slice(0, limit);
|
||||
return {
|
||||
status: "ok",
|
||||
credentials: selected.map((row) => toCredentialSummary(row, includePayload)),
|
||||
count: selected.length,
|
||||
};
|
||||
},
|
||||
});
|
||||
|
||||
export const cleanupLeaseEvents = internalMutation({
|
||||
args: {},
|
||||
handler: async (ctx) => {
|
||||
const cutoffMs = Date.now() - LEASE_EVENT_RETENTION_MS;
|
||||
const staleRows = await ctx.db
|
||||
.query("lease_events")
|
||||
.withIndex("by_occurredAtMs", (q) => q.lt("occurredAtMs", cutoffMs))
|
||||
.take(EVENT_RETENTION_BATCH_SIZE);
|
||||
|
||||
for (const row of staleRows) {
|
||||
await ctx.db.delete(row._id);
|
||||
}
|
||||
|
||||
if (staleRows.length === EVENT_RETENTION_BATCH_SIZE) {
|
||||
await ctx.scheduler.runAfter(0, internal.credentials.cleanupLeaseEvents, {});
|
||||
}
|
||||
|
||||
return {
|
||||
status: "ok",
|
||||
deleted: staleRows.length,
|
||||
retentionMs: LEASE_EVENT_RETENTION_MS,
|
||||
};
|
||||
},
|
||||
});
|
||||
|
||||
export const cleanupAdminEvents = internalMutation({
|
||||
args: {},
|
||||
handler: async (ctx) => {
|
||||
const cutoffMs = Date.now() - ADMIN_EVENT_RETENTION_MS;
|
||||
const staleRows = await ctx.db
|
||||
.query("admin_events")
|
||||
.withIndex("by_occurredAtMs", (q) => q.lt("occurredAtMs", cutoffMs))
|
||||
.take(EVENT_RETENTION_BATCH_SIZE);
|
||||
|
||||
for (const row of staleRows) {
|
||||
await ctx.db.delete(row._id);
|
||||
}
|
||||
|
||||
if (staleRows.length === EVENT_RETENTION_BATCH_SIZE) {
|
||||
await ctx.scheduler.runAfter(0, internal.credentials.cleanupAdminEvents, {});
|
||||
}
|
||||
|
||||
return {
|
||||
status: "ok",
|
||||
deleted: staleRows.length,
|
||||
retentionMs: ADMIN_EVENT_RETENTION_MS,
|
||||
};
|
||||
},
|
||||
});
|
||||
20
openclaw/qa/convex-credential-broker/convex/crons.ts
Normal file
20
openclaw/qa/convex-credential-broker/convex/crons.ts
Normal file
|
|
@ -0,0 +1,20 @@
|
|||
import { cronJobs } from "convex/server";
|
||||
import { internal } from "./_generated/api";
|
||||
|
||||
const crons = cronJobs();
|
||||
|
||||
crons.interval(
|
||||
"qa-credential-lease-event-retention",
|
||||
{ hours: 1 },
|
||||
internal.credentials.cleanupLeaseEvents,
|
||||
{},
|
||||
);
|
||||
|
||||
crons.interval(
|
||||
"qa-credential-admin-event-retention",
|
||||
{ hours: 1 },
|
||||
internal.credentials.cleanupAdminEvents,
|
||||
{},
|
||||
);
|
||||
|
||||
export default crons;
|
||||
457
openclaw/qa/convex-credential-broker/convex/http.ts
Normal file
457
openclaw/qa/convex-credential-broker/convex/http.ts
Normal file
|
|
@ -0,0 +1,457 @@
|
|||
import { httpRouter } from "convex/server";
|
||||
import { internal } from "./_generated/api";
|
||||
import type { Id } from "./_generated/dataModel";
|
||||
import { httpAction } from "./_generated/server";
|
||||
|
||||
type ActorRole = "ci" | "maintainer";
|
||||
|
||||
class BrokerHttpError extends Error {
|
||||
code: string;
|
||||
httpStatus: number;
|
||||
|
||||
constructor(httpStatus: number, code: string, message: string) {
|
||||
super(message);
|
||||
this.name = "BrokerHttpError";
|
||||
this.httpStatus = httpStatus;
|
||||
this.code = code;
|
||||
}
|
||||
}
|
||||
|
||||
function jsonResponse(status: number, payload: unknown) {
|
||||
return new Response(JSON.stringify(payload), {
|
||||
status,
|
||||
headers: {
|
||||
"content-type": "application/json; charset=utf-8",
|
||||
"cache-control": "no-store",
|
||||
},
|
||||
});
|
||||
}
|
||||
|
||||
function parseBearerToken(request: Request) {
|
||||
const header = request.headers.get("authorization")?.trim();
|
||||
if (!header) {
|
||||
return null;
|
||||
}
|
||||
const [scheme, token] = header.split(/\s+/u, 2);
|
||||
if (scheme?.toLowerCase() !== "bearer" || !token) {
|
||||
return null;
|
||||
}
|
||||
return token;
|
||||
}
|
||||
|
||||
function resolveAuthRole(token: string | null): ActorRole {
|
||||
if (!token) {
|
||||
throw new BrokerHttpError(
|
||||
401,
|
||||
"AUTH_REQUIRED",
|
||||
"Missing Authorization: Bearer <secret> header.",
|
||||
);
|
||||
}
|
||||
const maintainerSecret = process.env.OPENCLAW_QA_CONVEX_SECRET_MAINTAINER?.trim();
|
||||
const ciSecret = process.env.OPENCLAW_QA_CONVEX_SECRET_CI?.trim();
|
||||
|
||||
if (!maintainerSecret && !ciSecret) {
|
||||
throw new BrokerHttpError(
|
||||
500,
|
||||
"SERVER_MISCONFIGURED",
|
||||
"No Convex broker role secrets are configured on this deployment.",
|
||||
);
|
||||
}
|
||||
if (maintainerSecret && token === maintainerSecret) {
|
||||
return "maintainer";
|
||||
}
|
||||
if (ciSecret && token === ciSecret) {
|
||||
return "ci";
|
||||
}
|
||||
throw new BrokerHttpError(401, "AUTH_INVALID", "Credential broker secret is invalid.");
|
||||
}
|
||||
|
||||
function assertMaintainerAdminAuth(token: string | null) {
|
||||
if (!token) {
|
||||
throw new BrokerHttpError(
|
||||
401,
|
||||
"AUTH_REQUIRED",
|
||||
"Missing Authorization: Bearer <secret> header.",
|
||||
);
|
||||
}
|
||||
const maintainerSecret = process.env.OPENCLAW_QA_CONVEX_SECRET_MAINTAINER?.trim();
|
||||
if (!maintainerSecret) {
|
||||
throw new BrokerHttpError(
|
||||
500,
|
||||
"SERVER_MISCONFIGURED",
|
||||
"Admin endpoints require OPENCLAW_QA_CONVEX_SECRET_MAINTAINER on this deployment.",
|
||||
);
|
||||
}
|
||||
if (token === maintainerSecret) {
|
||||
return;
|
||||
}
|
||||
const ciSecret = process.env.OPENCLAW_QA_CONVEX_SECRET_CI?.trim();
|
||||
if (ciSecret && token === ciSecret) {
|
||||
throw new BrokerHttpError(
|
||||
403,
|
||||
"AUTH_ROLE_MISMATCH",
|
||||
"Admin endpoints require maintainer credentials.",
|
||||
);
|
||||
}
|
||||
throw new BrokerHttpError(401, "AUTH_INVALID", "Credential broker secret is invalid.");
|
||||
}
|
||||
|
||||
function asObject(value: unknown) {
|
||||
if (!value || typeof value !== "object" || Array.isArray(value)) {
|
||||
return null;
|
||||
}
|
||||
return value as Record<string, unknown>;
|
||||
}
|
||||
|
||||
async function parseJsonObject(request: Request) {
|
||||
let parsed: unknown;
|
||||
try {
|
||||
parsed = await request.json();
|
||||
} catch {
|
||||
throw new BrokerHttpError(400, "INVALID_JSON", "Request body must be valid JSON.");
|
||||
}
|
||||
const body = asObject(parsed);
|
||||
if (!body) {
|
||||
throw new BrokerHttpError(400, "INVALID_BODY", "Request body must be a JSON object.");
|
||||
}
|
||||
return body;
|
||||
}
|
||||
|
||||
function requireString(body: Record<string, unknown>, key: string) {
|
||||
const raw = body[key];
|
||||
if (typeof raw !== "string") {
|
||||
throw new BrokerHttpError(400, "INVALID_BODY", `Expected "${key}" to be a string.`);
|
||||
}
|
||||
const value = raw.trim();
|
||||
if (!value) {
|
||||
throw new BrokerHttpError(400, "INVALID_BODY", `Expected "${key}" to be non-empty.`);
|
||||
}
|
||||
return value;
|
||||
}
|
||||
|
||||
function optionalString(body: Record<string, unknown>, key: string) {
|
||||
if (!(key in body) || body[key] === undefined || body[key] === null) {
|
||||
return undefined;
|
||||
}
|
||||
const raw = body[key];
|
||||
if (typeof raw !== "string") {
|
||||
throw new BrokerHttpError(400, "INVALID_BODY", `Expected "${key}" to be a string.`);
|
||||
}
|
||||
const value = raw.trim();
|
||||
return value.length > 0 ? value : undefined;
|
||||
}
|
||||
|
||||
function requireObject(body: Record<string, unknown>, key: string) {
|
||||
const raw = body[key];
|
||||
const parsed = asObject(raw);
|
||||
if (!parsed) {
|
||||
throw new BrokerHttpError(400, "INVALID_BODY", `Expected "${key}" to be a JSON object.`);
|
||||
}
|
||||
return parsed;
|
||||
}
|
||||
|
||||
function optionalPositiveInteger(body: Record<string, unknown>, key: string) {
|
||||
if (!(key in body) || body[key] === undefined || body[key] === null) {
|
||||
return undefined;
|
||||
}
|
||||
const raw = body[key];
|
||||
if (typeof raw !== "number" || !Number.isFinite(raw) || !Number.isInteger(raw) || raw < 1) {
|
||||
throw new BrokerHttpError(400, "INVALID_BODY", `Expected "${key}" to be a positive integer.`);
|
||||
}
|
||||
return raw;
|
||||
}
|
||||
|
||||
function optionalBoolean(body: Record<string, unknown>, key: string) {
|
||||
if (!(key in body) || body[key] === undefined || body[key] === null) {
|
||||
return undefined;
|
||||
}
|
||||
if (typeof body[key] !== "boolean") {
|
||||
throw new BrokerHttpError(400, "INVALID_BODY", `Expected "${key}" to be a boolean.`);
|
||||
}
|
||||
return body[key];
|
||||
}
|
||||
|
||||
function optionalCredentialStatus(body: Record<string, unknown>, key: string) {
|
||||
const value = optionalString(body, key);
|
||||
if (!value) {
|
||||
return undefined;
|
||||
}
|
||||
if (value !== "active" && value !== "disabled") {
|
||||
throw new BrokerHttpError(
|
||||
400,
|
||||
"INVALID_BODY",
|
||||
`Expected "${key}" to be "active" or "disabled".`,
|
||||
);
|
||||
}
|
||||
return value;
|
||||
}
|
||||
|
||||
function optionalListStatus(body: Record<string, unknown>, key: string) {
|
||||
const value = optionalString(body, key);
|
||||
if (!value) {
|
||||
return undefined;
|
||||
}
|
||||
if (value !== "active" && value !== "disabled" && value !== "all") {
|
||||
throw new BrokerHttpError(
|
||||
400,
|
||||
"INVALID_BODY",
|
||||
`Expected "${key}" to be "active", "disabled", or "all".`,
|
||||
);
|
||||
}
|
||||
return value;
|
||||
}
|
||||
|
||||
function requirePayloadString(payload: Record<string, unknown>, key: string, kind: string): string {
|
||||
const raw = payload[key];
|
||||
if (typeof raw !== "string") {
|
||||
throw new BrokerHttpError(
|
||||
400,
|
||||
"INVALID_PAYLOAD",
|
||||
`Credential payload for kind "${kind}" must include "${key}" as a string.`,
|
||||
);
|
||||
}
|
||||
const value = raw.trim();
|
||||
if (!value) {
|
||||
throw new BrokerHttpError(
|
||||
400,
|
||||
"INVALID_PAYLOAD",
|
||||
`Credential payload for kind "${kind}" must include a non-empty "${key}" value.`,
|
||||
);
|
||||
}
|
||||
return value;
|
||||
}
|
||||
|
||||
function normalizeCredentialPayloadForKind(kind: string, payload: Record<string, unknown>) {
|
||||
if (kind !== "telegram") {
|
||||
return payload;
|
||||
}
|
||||
|
||||
const groupId = requirePayloadString(payload, "groupId", "telegram");
|
||||
if (!/^-?\d+$/u.test(groupId)) {
|
||||
throw new BrokerHttpError(
|
||||
400,
|
||||
"INVALID_PAYLOAD",
|
||||
'Credential payload for kind "telegram" must include a numeric "groupId" string.',
|
||||
);
|
||||
}
|
||||
|
||||
const driverToken = requirePayloadString(payload, "driverToken", "telegram");
|
||||
const sutToken = requirePayloadString(payload, "sutToken", "telegram");
|
||||
|
||||
return {
|
||||
groupId,
|
||||
driverToken,
|
||||
sutToken,
|
||||
} satisfies Record<string, unknown>;
|
||||
}
|
||||
|
||||
function parseActorRole(body: Record<string, unknown>) {
|
||||
const actorRole = requireString(body, "actorRole");
|
||||
if (actorRole !== "ci" && actorRole !== "maintainer") {
|
||||
throw new BrokerHttpError(
|
||||
400,
|
||||
"INVALID_ACTOR_ROLE",
|
||||
'Expected "actorRole" to be "maintainer" or "ci".',
|
||||
);
|
||||
}
|
||||
return actorRole as ActorRole;
|
||||
}
|
||||
|
||||
function assertRoleAllowed(tokenRole: ActorRole, requestedRole: ActorRole) {
|
||||
if (tokenRole !== requestedRole) {
|
||||
throw new BrokerHttpError(
|
||||
403,
|
||||
"AUTH_ROLE_MISMATCH",
|
||||
`Secret role "${tokenRole}" cannot be used as actorRole "${requestedRole}".`,
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
function normalizeCredentialId(raw: string) {
|
||||
// Convex Ids are opaque strings. We only enforce non-empty shape at HTTP boundary.
|
||||
return raw;
|
||||
}
|
||||
|
||||
function normalizeError(error: unknown) {
|
||||
if (error instanceof BrokerHttpError) {
|
||||
return {
|
||||
httpStatus: error.httpStatus,
|
||||
payload: {
|
||||
status: "error",
|
||||
code: error.code,
|
||||
message: error.message,
|
||||
},
|
||||
};
|
||||
}
|
||||
if (error instanceof Error) {
|
||||
return {
|
||||
httpStatus: 500,
|
||||
payload: {
|
||||
status: "error",
|
||||
code: "INTERNAL_ERROR",
|
||||
message: error.message || "Internal credential broker error.",
|
||||
},
|
||||
};
|
||||
}
|
||||
return {
|
||||
httpStatus: 500,
|
||||
payload: {
|
||||
status: "error",
|
||||
code: "INTERNAL_ERROR",
|
||||
message: "Internal credential broker error.",
|
||||
},
|
||||
};
|
||||
}
|
||||
|
||||
const http = httpRouter();
|
||||
|
||||
http.route({
|
||||
path: "/qa-credentials/v1/acquire",
|
||||
method: "POST",
|
||||
handler: httpAction(async (ctx, request) => {
|
||||
try {
|
||||
const tokenRole = resolveAuthRole(parseBearerToken(request));
|
||||
const body = await parseJsonObject(request);
|
||||
const actorRole = parseActorRole(body);
|
||||
assertRoleAllowed(tokenRole, actorRole);
|
||||
|
||||
const result = await ctx.runMutation(internal.credentials.acquireLease, {
|
||||
kind: requireString(body, "kind"),
|
||||
ownerId: requireString(body, "ownerId"),
|
||||
actorRole,
|
||||
leaseTtlMs: optionalPositiveInteger(body, "leaseTtlMs"),
|
||||
heartbeatIntervalMs: optionalPositiveInteger(body, "heartbeatIntervalMs"),
|
||||
});
|
||||
|
||||
return jsonResponse(200, result);
|
||||
} catch (error) {
|
||||
const normalized = normalizeError(error);
|
||||
return jsonResponse(normalized.httpStatus, normalized.payload);
|
||||
}
|
||||
}),
|
||||
});
|
||||
|
||||
http.route({
|
||||
path: "/qa-credentials/v1/heartbeat",
|
||||
method: "POST",
|
||||
handler: httpAction(async (ctx, request) => {
|
||||
try {
|
||||
const tokenRole = resolveAuthRole(parseBearerToken(request));
|
||||
const body = await parseJsonObject(request);
|
||||
const actorRole = parseActorRole(body);
|
||||
assertRoleAllowed(tokenRole, actorRole);
|
||||
|
||||
const result = await ctx.runMutation(internal.credentials.heartbeatLease, {
|
||||
kind: requireString(body, "kind"),
|
||||
ownerId: requireString(body, "ownerId"),
|
||||
actorRole,
|
||||
credentialId: normalizeCredentialId(
|
||||
requireString(body, "credentialId"),
|
||||
) as Id<"credential_sets">,
|
||||
leaseToken: requireString(body, "leaseToken"),
|
||||
leaseTtlMs: optionalPositiveInteger(body, "leaseTtlMs"),
|
||||
});
|
||||
|
||||
return jsonResponse(200, result);
|
||||
} catch (error) {
|
||||
const normalized = normalizeError(error);
|
||||
return jsonResponse(normalized.httpStatus, normalized.payload);
|
||||
}
|
||||
}),
|
||||
});
|
||||
|
||||
http.route({
|
||||
path: "/qa-credentials/v1/release",
|
||||
method: "POST",
|
||||
handler: httpAction(async (ctx, request) => {
|
||||
try {
|
||||
const tokenRole = resolveAuthRole(parseBearerToken(request));
|
||||
const body = await parseJsonObject(request);
|
||||
const actorRole = parseActorRole(body);
|
||||
assertRoleAllowed(tokenRole, actorRole);
|
||||
|
||||
const result = await ctx.runMutation(internal.credentials.releaseLease, {
|
||||
kind: requireString(body, "kind"),
|
||||
ownerId: requireString(body, "ownerId"),
|
||||
actorRole,
|
||||
credentialId: normalizeCredentialId(
|
||||
requireString(body, "credentialId"),
|
||||
) as Id<"credential_sets">,
|
||||
leaseToken: requireString(body, "leaseToken"),
|
||||
});
|
||||
|
||||
return jsonResponse(200, result);
|
||||
} catch (error) {
|
||||
const normalized = normalizeError(error);
|
||||
return jsonResponse(normalized.httpStatus, normalized.payload);
|
||||
}
|
||||
}),
|
||||
});
|
||||
|
||||
http.route({
|
||||
path: "/qa-credentials/v1/admin/add",
|
||||
method: "POST",
|
||||
handler: httpAction(async (ctx, request) => {
|
||||
try {
|
||||
assertMaintainerAdminAuth(parseBearerToken(request));
|
||||
const body = await parseJsonObject(request);
|
||||
const kind = requireString(body, "kind");
|
||||
const payload = normalizeCredentialPayloadForKind(kind, requireObject(body, "payload"));
|
||||
const result = await ctx.runMutation(internal.credentials.addCredentialSet, {
|
||||
kind,
|
||||
payload,
|
||||
note: optionalString(body, "note"),
|
||||
actorId: optionalString(body, "actorId"),
|
||||
status: optionalCredentialStatus(body, "status"),
|
||||
});
|
||||
return jsonResponse(200, result);
|
||||
} catch (error) {
|
||||
const normalized = normalizeError(error);
|
||||
return jsonResponse(normalized.httpStatus, normalized.payload);
|
||||
}
|
||||
}),
|
||||
});
|
||||
|
||||
http.route({
|
||||
path: "/qa-credentials/v1/admin/remove",
|
||||
method: "POST",
|
||||
handler: httpAction(async (ctx, request) => {
|
||||
try {
|
||||
assertMaintainerAdminAuth(parseBearerToken(request));
|
||||
const body = await parseJsonObject(request);
|
||||
const result = await ctx.runMutation(internal.credentials.disableCredentialSet, {
|
||||
credentialId: normalizeCredentialId(
|
||||
requireString(body, "credentialId"),
|
||||
) as Id<"credential_sets">,
|
||||
actorId: optionalString(body, "actorId"),
|
||||
});
|
||||
return jsonResponse(200, result);
|
||||
} catch (error) {
|
||||
const normalized = normalizeError(error);
|
||||
return jsonResponse(normalized.httpStatus, normalized.payload);
|
||||
}
|
||||
}),
|
||||
});
|
||||
|
||||
http.route({
|
||||
path: "/qa-credentials/v1/admin/list",
|
||||
method: "POST",
|
||||
handler: httpAction(async (ctx, request) => {
|
||||
try {
|
||||
assertMaintainerAdminAuth(parseBearerToken(request));
|
||||
const body = await parseJsonObject(request);
|
||||
const result = await ctx.runQuery(internal.credentials.listCredentialSets, {
|
||||
kind: optionalString(body, "kind"),
|
||||
status: optionalListStatus(body, "status"),
|
||||
includePayload: optionalBoolean(body, "includePayload"),
|
||||
limit: optionalPositiveInteger(body, "limit"),
|
||||
});
|
||||
return jsonResponse(200, result);
|
||||
} catch (error) {
|
||||
const normalized = normalizeError(error);
|
||||
return jsonResponse(normalized.httpStatus, normalized.payload);
|
||||
}
|
||||
}),
|
||||
});
|
||||
|
||||
export default http;
|
||||
63
openclaw/qa/convex-credential-broker/convex/schema.ts
Normal file
63
openclaw/qa/convex-credential-broker/convex/schema.ts
Normal file
|
|
@ -0,0 +1,63 @@
|
|||
import { defineSchema, defineTable } from "convex/server";
|
||||
import { v } from "convex/values";
|
||||
|
||||
const actorRole = v.union(v.literal("ci"), v.literal("maintainer"));
|
||||
const credentialStatus = v.union(v.literal("active"), v.literal("disabled"));
|
||||
const leaseEventType = v.union(
|
||||
v.literal("acquire"),
|
||||
v.literal("acquire_failed"),
|
||||
v.literal("release"),
|
||||
);
|
||||
const adminEventType = v.union(v.literal("add"), v.literal("disable"), v.literal("disable_failed"));
|
||||
|
||||
export default defineSchema({
|
||||
credential_sets: defineTable({
|
||||
kind: v.string(),
|
||||
status: credentialStatus,
|
||||
payload: v.any(),
|
||||
createdAtMs: v.number(),
|
||||
updatedAtMs: v.number(),
|
||||
lastLeasedAtMs: v.number(),
|
||||
note: v.optional(v.string()),
|
||||
lease: v.optional(
|
||||
v.object({
|
||||
ownerId: v.string(),
|
||||
actorRole,
|
||||
leaseToken: v.string(),
|
||||
acquiredAtMs: v.number(),
|
||||
heartbeatAtMs: v.number(),
|
||||
expiresAtMs: v.number(),
|
||||
}),
|
||||
),
|
||||
})
|
||||
.index("by_kind_status", ["kind", "status"])
|
||||
.index("by_kind_lastLeasedAtMs", ["kind", "lastLeasedAtMs"]),
|
||||
|
||||
lease_events: defineTable({
|
||||
kind: v.string(),
|
||||
eventType: leaseEventType,
|
||||
actorRole,
|
||||
ownerId: v.string(),
|
||||
occurredAtMs: v.number(),
|
||||
credentialId: v.optional(v.id("credential_sets")),
|
||||
code: v.optional(v.string()),
|
||||
message: v.optional(v.string()),
|
||||
})
|
||||
.index("by_occurredAtMs", ["occurredAtMs"])
|
||||
.index("by_kind_occurredAtMs", ["kind", "occurredAtMs"])
|
||||
.index("by_credential_occurredAtMs", ["credentialId", "occurredAtMs"]),
|
||||
|
||||
admin_events: defineTable({
|
||||
eventType: adminEventType,
|
||||
actorRole,
|
||||
actorId: v.string(),
|
||||
occurredAtMs: v.number(),
|
||||
credentialId: v.optional(v.id("credential_sets")),
|
||||
kind: v.optional(v.string()),
|
||||
code: v.optional(v.string()),
|
||||
message: v.optional(v.string()),
|
||||
})
|
||||
.index("by_occurredAtMs", ["occurredAtMs"])
|
||||
.index("by_kind_occurredAtMs", ["kind", "occurredAtMs"])
|
||||
.index("by_credential_occurredAtMs", ["credentialId", "occurredAtMs"]),
|
||||
});
|
||||
25
openclaw/qa/convex-credential-broker/convex/tsconfig.json
Normal file
25
openclaw/qa/convex-credential-broker/convex/tsconfig.json
Normal file
|
|
@ -0,0 +1,25 @@
|
|||
{
|
||||
/* This TypeScript project config describes the environment that
|
||||
* Convex functions run in and is used to typecheck them.
|
||||
* You can modify it, but some settings are required to use Convex.
|
||||
*/
|
||||
"compilerOptions": {
|
||||
/* These settings are not required by Convex and can be modified. */
|
||||
"allowJs": true,
|
||||
"strict": true,
|
||||
"moduleResolution": "Bundler",
|
||||
"jsx": "react-jsx",
|
||||
"skipLibCheck": true,
|
||||
"allowSyntheticDefaultImports": true,
|
||||
|
||||
/* These compiler options are required by Convex */
|
||||
"target": "ESNext",
|
||||
"lib": ["ES2023", "dom"],
|
||||
"forceConsistentCasingInFileNames": true,
|
||||
"module": "ESNext",
|
||||
"isolatedModules": true,
|
||||
"noEmit": true
|
||||
},
|
||||
"include": ["./**/*"],
|
||||
"exclude": ["./_generated"]
|
||||
}
|
||||
15
openclaw/qa/convex-credential-broker/package.json
Normal file
15
openclaw/qa/convex-credential-broker/package.json
Normal file
|
|
@ -0,0 +1,15 @@
|
|||
{
|
||||
"name": "@openclaw/qa-convex-credential-broker",
|
||||
"version": "0.1.0",
|
||||
"private": true,
|
||||
"description": "Convex HTTP credential lease broker for OpenClaw QA lab",
|
||||
"type": "module",
|
||||
"scripts": {
|
||||
"dashboard": "convex dashboard",
|
||||
"deploy": "convex deploy",
|
||||
"dev": "convex dev"
|
||||
},
|
||||
"dependencies": {
|
||||
"convex": "^1.35.1"
|
||||
}
|
||||
}
|
||||
132
openclaw/qa/frontier-harness-plan.md
Normal file
132
openclaw/qa/frontier-harness-plan.md
Normal file
|
|
@ -0,0 +1,132 @@
|
|||
# Frontier Harness Test Plan
|
||||
|
||||
Use this when tuning the harness on frontier models before the small-model pass.
|
||||
|
||||
## Goals
|
||||
|
||||
- verify tool-first behavior on short approval turns
|
||||
- verify model switching does not kill tool use
|
||||
- verify repo-reading / discovery still finishes with a concrete report
|
||||
- verify mutating work keeps replay-unsafety explicit under compaction pressure
|
||||
- collect manual notes on personality without letting style hide execution regressions
|
||||
|
||||
## Frontier subset
|
||||
|
||||
Run this subset first on every harness tweak:
|
||||
|
||||
- `approval-turn-tool-followthrough`
|
||||
- `model-switch-tool-continuity`
|
||||
- `source-docs-discovery-report`
|
||||
|
||||
Longer spot-check after that:
|
||||
|
||||
- `compaction-retry-mutating-tool`
|
||||
- `subagent-handoff`
|
||||
|
||||
## Baseline order
|
||||
|
||||
1. GPT first. Use this as the main tuning reference.
|
||||
2. Claude second. If Claude regresses alone, prefer an Anthropic overlay fix over a core prompt rewrite.
|
||||
3. Gemini third. Treat this as the operational-directness check.
|
||||
4. Only run the whole seed suite after the frontier subset is stable.
|
||||
|
||||
## Commands
|
||||
|
||||
GPT baseline:
|
||||
|
||||
```bash
|
||||
pnpm openclaw qa suite \
|
||||
--provider-mode live-frontier \
|
||||
--model openai/gpt-5.4 \
|
||||
--alt-model openai/gpt-5.4 \
|
||||
--fast \
|
||||
--scenario approval-turn-tool-followthrough \
|
||||
--scenario model-switch-tool-continuity \
|
||||
--scenario source-docs-discovery-report
|
||||
```
|
||||
|
||||
Claude sweep:
|
||||
|
||||
```bash
|
||||
pnpm openclaw qa suite \
|
||||
--provider-mode live-frontier \
|
||||
--model anthropic/claude-sonnet-4-6 \
|
||||
--alt-model anthropic/claude-opus-4-6 \
|
||||
--scenario approval-turn-tool-followthrough \
|
||||
--scenario model-switch-tool-continuity \
|
||||
--scenario source-docs-discovery-report
|
||||
```
|
||||
|
||||
Gemini sweep:
|
||||
|
||||
```bash
|
||||
pnpm openclaw qa suite \
|
||||
--provider-mode live-frontier \
|
||||
--model <google-pro-model-ref> \
|
||||
--alt-model <google-pro-model-ref> \
|
||||
--scenario approval-turn-tool-followthrough \
|
||||
--scenario model-switch-tool-continuity \
|
||||
--scenario source-docs-discovery-report
|
||||
```
|
||||
|
||||
Use the QA Lab runner catalog or `openclaw models list --all` to pick the current Google Pro ref.
|
||||
|
||||
## Tuning loop
|
||||
|
||||
1. Run the GPT subset and save the report path.
|
||||
2. Patch one harness idea at a time.
|
||||
3. Rerun the same GPT subset immediately.
|
||||
4. If GPT improves, run the Claude subset.
|
||||
5. If Claude is clean, run the Gemini subset.
|
||||
6. If only one family regresses, fix the provider overlay before touching the shared prompt again.
|
||||
|
||||
## What to score
|
||||
|
||||
- tool commitment after `ok do it`
|
||||
- empty-promise rate
|
||||
- tool continuity after model switch
|
||||
- discovery report completeness and specificity
|
||||
- replay-safety truth after a mutating write
|
||||
- scope drift: unrelated scenario updates, grand wrap-ups, or invented completion tallies
|
||||
- latency / obvious stall behavior
|
||||
- token cost notes if a change makes the prompt materially heavier
|
||||
|
||||
## Manual personality lane
|
||||
|
||||
Run this after the executable subset, not before:
|
||||
|
||||
```text
|
||||
read QA_KICKOFF_TASK.md, tell me what feels half-baked about this qa mission, and keep it to two short sentences
|
||||
```
|
||||
|
||||
GPT manual lane:
|
||||
|
||||
```bash
|
||||
pnpm openclaw qa manual \
|
||||
--provider-mode live-frontier \
|
||||
--model openai/gpt-5.4 \
|
||||
--alt-model openai/gpt-5.4 \
|
||||
--fast \
|
||||
--message "read QA_KICKOFF_TASK.md, tell me what feels half-baked about this qa mission, and keep it to two short sentences"
|
||||
```
|
||||
|
||||
Claude manual lane:
|
||||
|
||||
```bash
|
||||
pnpm openclaw qa manual \
|
||||
--provider-mode live-frontier \
|
||||
--model anthropic/claude-sonnet-4-6 \
|
||||
--alt-model anthropic/claude-opus-4-6 \
|
||||
--message "read QA_KICKOFF_TASK.md, tell me what feels half-baked about this qa mission, and keep it to two short sentences"
|
||||
```
|
||||
|
||||
Score it on:
|
||||
|
||||
- did it read first
|
||||
- did it say something specific instead of generic fluff
|
||||
- did the agent still sound like itself while doing useful work
|
||||
- did it stay on the scoped ask instead of widening into a suite recap or fake completion claim
|
||||
|
||||
## Deferred
|
||||
|
||||
- deterministic mock compaction triggering is still deferred; the current replay-safety lane is a live-frontier-first executable scenario
|
||||
151
openclaw/qa/new-scenarios-2026-04.md
Normal file
151
openclaw/qa/new-scenarios-2026-04.md
Normal file
|
|
@ -0,0 +1,151 @@
|
|||
# QA Scenario Expansion - Round 2
|
||||
|
||||
Ten repo-grounded candidate scenarios to add after the current seed suite.
|
||||
|
||||
## 1. On-demand memory tools in channel context
|
||||
|
||||
- Goal: verify the agent uses `memory_search` plus `memory_get` instead of bluffing when a channel message asks about prior notes.
|
||||
- Flow:
|
||||
- Seed `MEMORY.md` or `memory/*.md` with a fact not present in the current transcript.
|
||||
- Ask in a channel thread for that fact.
|
||||
- Verify tool usage and final answer accuracy.
|
||||
- Pass:
|
||||
- `memory_search` runs first.
|
||||
- `memory_get` narrows to the right lines.
|
||||
- Final answer cites the remembered fact correctly without cross-session leakage.
|
||||
- Docs: `docs/concepts/memory.md`, `docs/concepts/memory-search.md`
|
||||
- Code: `extensions/memory-core/src/tools.ts`, `extensions/memory-core/src/prompt-section.ts`
|
||||
|
||||
## 2. Memory failure fallback
|
||||
|
||||
- Goal: verify memory failure is graceful when embeddings/search are unavailable.
|
||||
- Flow:
|
||||
- Disable or break the embedding-backed memory path.
|
||||
- Ask for prior-note recall.
|
||||
- Verify the agent surfaces uncertainty and next action instead of hallucinating.
|
||||
- Pass:
|
||||
- Tool failure does not crash the run.
|
||||
- Agent says it checked and could not confirm.
|
||||
- Report includes the remediation hint.
|
||||
- Docs: `docs/concepts/memory.md`, `docs/help/faq.md`
|
||||
- Code: `extensions/memory-core/src/tools.shared.ts`, `extensions/memory-core/src/tools.citations.test.ts`
|
||||
|
||||
## 3. Model switch with tool continuity
|
||||
|
||||
- Goal: verify model switching preserves session context and tool availability, not just plain text continuity.
|
||||
- Flow:
|
||||
- Start on one model.
|
||||
- Switch to another configured model.
|
||||
- Ask for a tool-using follow-up such as file read or memory lookup.
|
||||
- Pass:
|
||||
- Switch is reflected in runtime state.
|
||||
- Tool call still succeeds after the switch.
|
||||
- Final answer keeps prior context.
|
||||
- Docs: `docs/help/testing.md`, `docs/concepts/model-failover.md`
|
||||
- Code: `extensions/qa-lab/src/suite.ts`, `docs/web/webchat.md`
|
||||
|
||||
## 4. MCP-backed recall via QMD/mcporter
|
||||
|
||||
- Goal: verify an MCP-backed tool path works end to end, not just core tools.
|
||||
- Flow:
|
||||
- Enable `memory.qmd.mcporter`.
|
||||
- Ask for recall that should route through the QMD MCP bridge.
|
||||
- Verify response and captured MCP execution path.
|
||||
- Pass:
|
||||
- MCP-backed search path is used.
|
||||
- Returned snippet matches the right note.
|
||||
- Failure mode is explicit if the daemon/tool is missing.
|
||||
- Docs: `docs/gateway/secrets.md`, `docs/concepts/memory-qmd.md`
|
||||
- Code: `extensions/memory-core/src/memory/qmd-manager.ts`, `extensions/memory-core/src/memory/qmd-manager.test.ts`
|
||||
|
||||
## 5. Skill visibility and invocation
|
||||
|
||||
- Goal: verify the agent sees a workspace/project skill and actually uses it.
|
||||
- Flow:
|
||||
- Add a simple workspace or `.agents` skill.
|
||||
- Confirm skill visibility through runtime inventory.
|
||||
- Ask for a task that should trigger the skill.
|
||||
- Pass:
|
||||
- Skill appears in `skills.status`.
|
||||
- Agent invocation reflects the installed skill instructions.
|
||||
- Per-agent allowlist behavior is respected.
|
||||
- Docs: `docs/tools/skills.md`, `docs/gateway/protocol.md`, `docs/gateway/configuration.md`
|
||||
- Code: `.agents/skills/openclaw-qa-testing/SKILL.md`, `docs/gateway/protocol.md`
|
||||
|
||||
## 6. Skill install and hot availability
|
||||
|
||||
- Goal: verify a newly installed skill becomes usable without a broken intermediate state.
|
||||
- Flow:
|
||||
- Install a ClawHub or gateway-managed skill.
|
||||
- Re-check skill inventory.
|
||||
- Ask the agent to perform the skill-backed task.
|
||||
- Pass:
|
||||
- Install succeeds.
|
||||
- `skills.status` or `skills.bins` reflects the new skill.
|
||||
- Agent can use the skill immediately or after the expected reload path.
|
||||
- Docs: `docs/tools/skills.md`, `docs/cli/skills.md`, `docs/gateway/protocol.md`
|
||||
- Code: `docs/gateway/protocol.md`, `docs/tools/skills.md`
|
||||
|
||||
## 7. Native image generation
|
||||
|
||||
- Goal: verify `image_generate` appears only when configured and returns a real attachment/artifact.
|
||||
- Flow:
|
||||
- Configure `agents.defaults.imageGenerationModel.primary`.
|
||||
- Ask for a simple generated image.
|
||||
- Verify generated media is returned in the reply path.
|
||||
- Pass:
|
||||
- `image_generate` is in the effective tool set.
|
||||
- Generation succeeds with the configured provider/model.
|
||||
- Output is attached and the agent summarizes what it created.
|
||||
- Docs: `docs/tools/image-generation.md`, `docs/providers/openai.md`
|
||||
- Code: `src/agents/openclaw-tools.image-generation.test.ts`, `src/image-generation/runtime.ts`
|
||||
|
||||
## 8. Config patch skill disable
|
||||
|
||||
- Goal: verify `config.patch` can disable a workspace skill and the restarted gateway exposes the disabled state cleanly.
|
||||
- Flow:
|
||||
- Add a workspace skill and verify it is eligible.
|
||||
- Use `config.patch` to disable that skill.
|
||||
- Wait for the gateway restart and read `skills.status` again.
|
||||
- Pass:
|
||||
- Patch succeeds.
|
||||
- Gateway restarts cleanly.
|
||||
- The skill flips from eligible to disabled.
|
||||
- Docs: `docs/gateway/configuration.md`, `docs/gateway/protocol.md`
|
||||
- Code: `docs/gateway/configuration.md`, `docs/web/control-ui.md`
|
||||
|
||||
## 9. Restart-required config apply with wake-up
|
||||
|
||||
- Goal: verify a restart-required config change restarts cleanly and wakes the session back up.
|
||||
- Flow:
|
||||
- Use `config.apply` or `update.run` on a restart-required surface.
|
||||
- Provide `sessionKey` so the operator gets the post-restart ping.
|
||||
- Resume the task after restart.
|
||||
- Pass:
|
||||
- Restart happens once.
|
||||
- Session wake-up ping arrives.
|
||||
- Agent continues in the same logical workflow after restart.
|
||||
- Docs: `docs/gateway/configuration.md`, `docs/web/control-ui.md`
|
||||
- Code: `docs/gateway/configuration.md`, `docs/gateway/protocol.md`
|
||||
|
||||
## 10. Runtime inventory drift check
|
||||
|
||||
- Goal: verify the reported tool and skill inventory matches what the agent can really use after config/plugin changes.
|
||||
- Flow:
|
||||
- Read `tools.effective` and `skills.status`.
|
||||
- Ask the agent to use one enabled thing and one disabled thing.
|
||||
- Compare actual behavior vs reported inventory.
|
||||
- Pass:
|
||||
- Enabled item is callable.
|
||||
- Disabled item is absent or blocked for the right reason.
|
||||
- Inventory and runtime behavior stay in sync.
|
||||
- Docs: `docs/gateway/protocol.md`, `docs/web/webchat.md`
|
||||
- Code: `docs/gateway/protocol.md`, `docs/web/control-ui.md`
|
||||
|
||||
## Best next additions to the executable suite
|
||||
|
||||
If we only promote three right away:
|
||||
|
||||
1. On-demand memory tools in channel context
|
||||
2. Native image generation
|
||||
3. Config patch skill disable
|
||||
8
openclaw/qa/scenarios.md
Normal file
8
openclaw/qa/scenarios.md
Normal file
|
|
@ -0,0 +1,8 @@
|
|||
# OpenClaw QA Scenarios
|
||||
|
||||
Canonical scenario source now lives in:
|
||||
|
||||
- `qa/scenarios/index.md`
|
||||
- `qa/scenarios/<theme>/*.md`
|
||||
|
||||
Each QA scenario has its own markdown file.
|
||||
|
|
@ -0,0 +1,145 @@
|
|||
# Instruction followthrough repo contract
|
||||
|
||||
```yaml qa-scenario
|
||||
id: instruction-followthrough-repo-contract
|
||||
title: Instruction followthrough repo contract
|
||||
surface: repo-contract
|
||||
coverage:
|
||||
primary:
|
||||
- agents.instructions
|
||||
secondary:
|
||||
- runtime.first-action
|
||||
objective: Verify the agent reads repo instruction files first, follows the required tool order, and completes the first feasible action instead of stopping at a plan.
|
||||
successCriteria:
|
||||
- Agent reads the seeded instruction files before writing the requested artifact.
|
||||
- Agent writes the requested artifact in the same run instead of returning only a plan.
|
||||
- Agent does not ask for permission before the first feasible action.
|
||||
- Final reply makes the completed read/write sequence explicit.
|
||||
docsRefs:
|
||||
- docs/help/testing.md
|
||||
- docs/channels/qa-channel.md
|
||||
codeRefs:
|
||||
- src/agents/system-prompt.ts
|
||||
- src/agents/pi-embedded-runner/run/incomplete-turn.ts
|
||||
- extensions/qa-lab/src/mock-openai-server.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify the agent reads repo instructions first, then completes the first bounded followthrough task without stalling.
|
||||
config:
|
||||
workspaceFiles:
|
||||
AGENT.md: |-
|
||||
# Repo contract
|
||||
|
||||
Step order:
|
||||
1. Read AGENT.md.
|
||||
2. Read SOUL.md.
|
||||
3. Read FOLLOWTHROUGH_INPUT.md.
|
||||
4. Write ./repo-contract-summary.txt.
|
||||
5. Reply with three labeled lines exactly once: Read, Wrote, Status.
|
||||
|
||||
Do not stop after planning.
|
||||
Do not ask for permission before the first feasible action.
|
||||
SOUL.md: |-
|
||||
# Execution style
|
||||
|
||||
Stay brief, honest, and action-first.
|
||||
If the next tool action is feasible, do it before replying.
|
||||
FOLLOWTHROUGH_INPUT.md: |-
|
||||
Mission: prove you followed the repo contract.
|
||||
Evidence path: AGENT.md -> SOUL.md -> FOLLOWTHROUGH_INPUT.md -> repo-contract-summary.txt
|
||||
prompt: |-
|
||||
Repo contract followthrough check. Read AGENT.md, SOUL.md, and FOLLOWTHROUGH_INPUT.md first.
|
||||
Then follow the repo contract exactly, write ./repo-contract-summary.txt, and reply with
|
||||
three labeled lines: Read, Wrote, Status.
|
||||
Do not stop after planning and do not ask for permission before the first feasible action.
|
||||
expectedReplyAll:
|
||||
- "read:"
|
||||
- "wrote:"
|
||||
- "status:"
|
||||
expectedArtifactAll:
|
||||
- "repo contract"
|
||||
expectedArtifactAny:
|
||||
- "evidence path"
|
||||
- "agent.md"
|
||||
- "followthrough"
|
||||
forbiddenNeedles:
|
||||
- need permission
|
||||
- need your approval
|
||||
- can you approve
|
||||
- i would
|
||||
- i can
|
||||
- next i would
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: follows repo instructions instead of stopping at a plan
|
||||
actions:
|
||||
- call: reset
|
||||
- forEach:
|
||||
items:
|
||||
expr: "Object.entries(config.workspaceFiles ?? {})"
|
||||
item: workspaceFile
|
||||
actions:
|
||||
- call: fs.writeFile
|
||||
args:
|
||||
- expr: "path.join(env.gateway.workspaceDir, String(workspaceFile[0]))"
|
||||
- expr: "`${String(workspaceFile[1] ?? '').trimEnd()}\\n`"
|
||||
- utf8
|
||||
- set: artifactPath
|
||||
value:
|
||||
expr: "path.join(env.gateway.workspaceDir, 'repo-contract-summary.txt')"
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:repo-contract
|
||||
message:
|
||||
expr: config.prompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 40000)
|
||||
- call: waitForCondition
|
||||
saveAs: artifact
|
||||
args:
|
||||
- lambda:
|
||||
async: true
|
||||
expr: "(() => { const normalize = (value) => normalizeLowercaseStringOrEmpty(value); const matches = (value) => { const normalized = normalize(value); return normalized && config.expectedArtifactAll.every((needle) => normalized.includes(normalize(needle))) && config.expectedArtifactAny.some((needle) => normalized.includes(normalize(needle))); }; return fs.readFile(artifactPath, 'utf8').then((value) => matches(value) ? value : undefined).catch(() => undefined); })()"
|
||||
- expr: liveTurnTimeoutMs(env, 30000)
|
||||
- expr: "env.providerMode === 'mock-openai' ? 100 : 250"
|
||||
- set: normalizedArtifact
|
||||
value:
|
||||
expr: "normalizeLowercaseStringOrEmpty(artifact)"
|
||||
- assert:
|
||||
expr: "config.expectedArtifactAll.every((needle) => normalizedArtifact.includes(normalizeLowercaseStringOrEmpty(needle))) && config.expectedArtifactAny.some((needle) => normalizedArtifact.includes(normalizeLowercaseStringOrEmpty(needle)))"
|
||||
message:
|
||||
expr: "`repo contract artifact missing expected followthrough signals: ${artifact}`"
|
||||
- set: expectedReplyAll
|
||||
value:
|
||||
expr: config.expectedReplyAll.map(normalizeLowercaseStringOrEmpty)
|
||||
- call: waitForCondition
|
||||
saveAs: outbound
|
||||
args:
|
||||
- lambda:
|
||||
expr: "state.getSnapshot().messages.filter((candidate) => candidate.direction === 'outbound' && candidate.conversation.id === 'qa-operator' && expectedReplyAll.every((needle) => normalizeLowercaseStringOrEmpty(candidate.text).includes(needle))).at(-1)"
|
||||
- expr: liveTurnTimeoutMs(env, 30000)
|
||||
- expr: "env.providerMode === 'mock-openai' ? 100 : 250"
|
||||
- assert:
|
||||
expr: "!config.forbiddenNeedles.some((needle) => normalizeLowercaseStringOrEmpty(outbound.text).includes(needle))"
|
||||
message:
|
||||
expr: "`repo contract followthrough bounced for permission or stalled: ${outbound.text}`"
|
||||
- set: followthroughDebugRequests
|
||||
value:
|
||||
expr: "env.mock ? [...(await fetchJson(`${env.mock.baseUrl}/debug/requests`))].filter((request) => /repo contract followthrough check/i.test(String(request.allInputText ?? ''))) : []"
|
||||
- assert:
|
||||
expr: "!env.mock || followthroughDebugRequests.filter((request) => request.plannedToolName === 'read').length >= 3"
|
||||
message:
|
||||
expr: "`expected three read tool calls before write, saw plannedToolNames=${JSON.stringify(followthroughDebugRequests.map((request) => request.plannedToolName ?? null))}`"
|
||||
- assert:
|
||||
expr: "!env.mock || followthroughDebugRequests.some((request) => request.plannedToolName === 'write')"
|
||||
message:
|
||||
expr: "`expected write tool call during repo contract followthrough, saw plannedToolNames=${JSON.stringify(followthroughDebugRequests.map((request) => request.plannedToolName ?? null))}`"
|
||||
- assert:
|
||||
expr: "!env.mock || (() => { const readIndices = followthroughDebugRequests.map((r, i) => r.plannedToolName === 'read' ? i : -1).filter(i => i >= 0); const firstWrite = followthroughDebugRequests.findIndex((r) => r.plannedToolName === 'write'); return readIndices.length >= 3 && firstWrite >= 0 && readIndices[2] < firstWrite; })()"
|
||||
message:
|
||||
expr: "`expected all 3 reads before any write during repo contract followthrough, saw plannedToolNames=${JSON.stringify(followthroughDebugRequests.map((request) => request.plannedToolName ?? null))}`"
|
||||
detailsExpr: outbound.text
|
||||
```
|
||||
170
openclaw/qa/scenarios/agents/subagent-fanout-synthesis.md
Normal file
170
openclaw/qa/scenarios/agents/subagent-fanout-synthesis.md
Normal file
|
|
@ -0,0 +1,170 @@
|
|||
# Subagent fanout synthesis
|
||||
|
||||
```yaml qa-scenario
|
||||
id: subagent-fanout-synthesis
|
||||
title: Subagent fanout synthesis
|
||||
surface: subagents
|
||||
coverage:
|
||||
primary:
|
||||
- agents.subagents
|
||||
secondary:
|
||||
- agents.synthesis
|
||||
objective: Verify the agent can delegate multiple bounded subagent tasks and fold both results back into one parent reply.
|
||||
successCriteria:
|
||||
- Parent flow launches at least two bounded subagent tasks.
|
||||
- Both delegated results are acknowledged in the main flow.
|
||||
- Final answer synthesizes both worker outputs in one reply.
|
||||
docsRefs:
|
||||
- docs/tools/subagents.md
|
||||
- docs/help/testing.md
|
||||
codeRefs:
|
||||
- src/agents/subagent-spawn.ts
|
||||
- src/agents/system-prompt.ts
|
||||
- extensions/qa-lab/src/suite.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify the agent can delegate multiple bounded subagent tasks and fold both results back into one parent reply.
|
||||
config:
|
||||
prompt: |-
|
||||
Subagent fanout synthesis check: delegate exactly two bounded subagents sequentially.
|
||||
Subagent 1: verify that `HEARTBEAT.md` exists and report `ok` if it does.
|
||||
Subagent 2: verify that `repo/qa/scenarios/agents/subagent-fanout-synthesis.md` exists and report `ok` if it does.
|
||||
Wait for both subagents to finish.
|
||||
Then reply with exactly these two lines and nothing else:
|
||||
subagent-1: ok
|
||||
subagent-2: ok
|
||||
Do not use ACP.
|
||||
expectedReplyAny:
|
||||
- "subagent-1: ok"
|
||||
- "subagent-2: ok"
|
||||
expectedReplyGroups:
|
||||
- - alpha-ok
|
||||
- subagent_one_ok
|
||||
- subagent one ok
|
||||
- "subagent-1: ok"
|
||||
- - beta-ok
|
||||
- subagent_two_ok
|
||||
- subagent two ok
|
||||
- "subagent-2: ok"
|
||||
expectedChildLabels:
|
||||
- qa-fanout-alpha
|
||||
- qa-fanout-beta
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: spawns sequential workers and folds both results back into the parent reply
|
||||
actions:
|
||||
- set: attempts
|
||||
value:
|
||||
expr: "env.providerMode === 'mock-openai' ? 1 : 2"
|
||||
- set: lastError
|
||||
value: null
|
||||
- forEach:
|
||||
items:
|
||||
expr: "Array.from({ length: attempts }, (_, index) => index + 1)"
|
||||
item: attempt
|
||||
actions:
|
||||
- if:
|
||||
expr: "lastError === '__done__'"
|
||||
then:
|
||||
- set: skippedAttempt
|
||||
value:
|
||||
expr: attempt
|
||||
else:
|
||||
- try:
|
||||
actions:
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- 120000
|
||||
- call: reset
|
||||
- set: sessionKey
|
||||
value:
|
||||
expr: "`agent:qa:fanout:${attempt}:${randomUUID().slice(0, 8)}`"
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey:
|
||||
ref: sessionKey
|
||||
message:
|
||||
expr: config.prompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 90000)
|
||||
- call: waitForCondition
|
||||
saveAs: outbound
|
||||
args:
|
||||
- lambda:
|
||||
expr: "state.getSnapshot().messages.filter((message) => message.direction === 'outbound' && message.conversation.id === 'qa-operator' && config.expectedReplyGroups.every((group) => group.some((needle) => normalizeLowercaseStringOrEmpty(message.text ?? '').includes(needle)))).at(-1)"
|
||||
- expr: liveTurnTimeoutMs(env, 60000)
|
||||
- expr: "env.providerMode === 'mock-openai' ? 100 : 250"
|
||||
- if:
|
||||
expr: "Boolean(env.mock)"
|
||||
then:
|
||||
- call: readRawQaSessionStore
|
||||
saveAs: store
|
||||
args:
|
||||
- ref: env
|
||||
- set: childRows
|
||||
value:
|
||||
expr: "Object.values(store).filter((entry) => entry.spawnedBy === sessionKey)"
|
||||
- set: sawAlpha
|
||||
value:
|
||||
expr: "childRows.some((entry) => entry.label === config.expectedChildLabels[0])"
|
||||
- set: sawBeta
|
||||
value:
|
||||
expr: "childRows.some((entry) => entry.label === config.expectedChildLabels[1])"
|
||||
- assert:
|
||||
expr: "sawAlpha && sawBeta"
|
||||
message:
|
||||
expr: "`fanout child sessions missing (alpha=${String(sawAlpha)} beta=${String(sawBeta)})`"
|
||||
# Tool-call assertion (criterion 2 of the
|
||||
# parity completion gate in #64227): the
|
||||
# scenario must have actually invoked
|
||||
# `sessions_spawn` at least twice with
|
||||
# distinct labels, not just ended up with
|
||||
# two rows in the session store through
|
||||
# prose trickery. The session store alone
|
||||
# can be populated by other flows or by a
|
||||
# model that fabricates "delegation"
|
||||
# narration. `plannedToolName` on the
|
||||
# mock's `/debug/requests` log is the
|
||||
# tool-call ground truth: two recorded
|
||||
# sessions_spawn requests with distinct
|
||||
# labels means the model really dispatched
|
||||
# both subagents.
|
||||
- set: fanoutSpawnRequests
|
||||
value:
|
||||
expr: "[...(await fetchJson(`${env.mock.baseUrl}/debug/requests`))].filter((request) => request.plannedToolName === 'sessions_spawn' && /subagent fanout synthesis check/i.test(String(request.allInputText ?? '')))"
|
||||
- assert:
|
||||
expr: "fanoutSpawnRequests.length >= 2"
|
||||
message:
|
||||
expr: "`expected at least two sessions_spawn tool calls during subagent fanout scenario, saw ${fanoutSpawnRequests.length}`"
|
||||
- set: details
|
||||
value:
|
||||
expr: "outbound.text"
|
||||
- set: lastError
|
||||
value: __done__
|
||||
catchAs: attemptError
|
||||
catch:
|
||||
- set: lastError
|
||||
value:
|
||||
ref: attemptError
|
||||
- if:
|
||||
expr: "attempt < attempts"
|
||||
then:
|
||||
- try:
|
||||
actions:
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- 120000
|
||||
catch:
|
||||
- set: ignoredRetryWait
|
||||
value: true
|
||||
- assert:
|
||||
expr: "lastError === '__done__'"
|
||||
message:
|
||||
expr: "lastError instanceof Error ? formatErrorMessage(lastError) : String(lastError ?? 'fanout retry exhausted')"
|
||||
detailsExpr: "details"
|
||||
```
|
||||
73
openclaw/qa/scenarios/agents/subagent-handoff.md
Normal file
73
openclaw/qa/scenarios/agents/subagent-handoff.md
Normal file
|
|
@ -0,0 +1,73 @@
|
|||
# Subagent handoff
|
||||
|
||||
```yaml qa-scenario
|
||||
id: subagent-handoff
|
||||
title: Subagent handoff
|
||||
surface: subagents
|
||||
coverage:
|
||||
primary:
|
||||
- agents.subagents
|
||||
objective: Verify the agent can delegate a bounded task to a subagent and fold the result back into the main thread.
|
||||
successCriteria:
|
||||
- Agent launches a bounded subagent task.
|
||||
- Subagent result is acknowledged in the main flow.
|
||||
- Final answer attributes delegated work clearly.
|
||||
docsRefs:
|
||||
- docs/tools/subagents.md
|
||||
- docs/help/testing.md
|
||||
codeRefs:
|
||||
- src/agents/system-prompt.ts
|
||||
- extensions/qa-lab/src/report.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify the agent can delegate a bounded task to a subagent and fold the result back into the main thread.
|
||||
config:
|
||||
prompt: "Delegate one bounded QA task to a subagent. Wait for the subagent to finish. Then reply with three labeled sections exactly once: Delegated task, Result, Evidence. Include the child result itself, not 'waiting'."
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: delegates a bounded task and reports the result
|
||||
actions:
|
||||
- call: reset
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:subagent
|
||||
message:
|
||||
expr: config.prompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 90000)
|
||||
- call: waitForCondition
|
||||
saveAs: outbound
|
||||
args:
|
||||
- lambda:
|
||||
expr: "state.getSnapshot().messages.filter((candidate) => candidate.direction === 'outbound' && candidate.conversation.id === 'qa-operator' && (() => { const lower = normalizeLowercaseStringOrEmpty(candidate.text); return lower.includes('delegated task') && lower.includes('result') && lower.includes('evidence') && !lower.includes('waiting'); })()).at(-1)"
|
||||
- expr: liveTurnTimeoutMs(env, 45000)
|
||||
- expr: "env.providerMode === 'mock-openai' ? 100 : 250"
|
||||
- assert:
|
||||
expr: "!['failed to delegate','could not delegate','subagent unavailable'].some((needle) => normalizeLowercaseStringOrEmpty(outbound.text).includes(needle))"
|
||||
message:
|
||||
expr: "`subagent handoff reported failure: ${outbound.text}`"
|
||||
# Parity gate criterion 2 (no fake progress / fake tool completion):
|
||||
# require an actual sessions_spawn tool call. Without this, a model
|
||||
# could produce the three labeled sections ("Delegated task", "Result",
|
||||
# "Evidence") as free-form prose without ever delegating to a real
|
||||
# subagent. The assertion is pinned to THIS scenario by matching the
|
||||
# scenario-unique prompt substring "Delegate one bounded QA task"
|
||||
# (not a broad /delegate|subagent/ regex) so the earlier
|
||||
# subagent-fanout-synthesis scenario — which also contains "delegate"
|
||||
# and produces its own pre-tool sessions_spawn request — cannot
|
||||
# satisfy the assertion here. The match is also constrained to
|
||||
# pre-tool requests (no toolOutput) because the mock only plans
|
||||
# sessions_spawn on requests with no toolOutput; the follow-up
|
||||
# request after the tool runs has plannedToolName unset.
|
||||
- set: subagentDebugRequests
|
||||
value:
|
||||
expr: "env.mock ? [...(await fetchJson(`${env.mock.baseUrl}/debug/requests`))] : []"
|
||||
- assert:
|
||||
expr: "!env.mock || subagentDebugRequests.some((request) => !request.toolOutput && /delegate one bounded qa task/i.test(String(request.allInputText ?? '')) && request.plannedToolName === 'sessions_spawn')"
|
||||
message:
|
||||
expr: "`expected sessions_spawn tool call during subagent handoff scenario, saw plannedToolNames=${JSON.stringify(subagentDebugRequests.map((request) => request.plannedToolName ?? null))}`"
|
||||
detailsExpr: outbound.text
|
||||
```
|
||||
84
openclaw/qa/scenarios/channels/channel-chat-baseline.md
Normal file
84
openclaw/qa/scenarios/channels/channel-chat-baseline.md
Normal file
|
|
@ -0,0 +1,84 @@
|
|||
# Channel baseline conversation
|
||||
|
||||
```yaml qa-scenario
|
||||
id: channel-chat-baseline
|
||||
title: Channel baseline conversation
|
||||
surface: channel
|
||||
coverage:
|
||||
primary:
|
||||
- channels.group-messages
|
||||
secondary:
|
||||
- channels.qa-channel
|
||||
objective: Verify the QA agent can respond correctly in a shared channel and respect mention-driven group semantics.
|
||||
successCriteria:
|
||||
- Agent replies in the shared channel transcript.
|
||||
- Agent keeps the conversation scoped to the channel.
|
||||
- Agent respects mention-driven group routing semantics.
|
||||
docsRefs:
|
||||
- docs/channels/group-messages.md
|
||||
- docs/channels/qa-channel.md
|
||||
codeRefs:
|
||||
- extensions/qa-channel/src/inbound.ts
|
||||
- extensions/qa-lab/src/bus-state.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify the QA agent can respond correctly in a shared channel and respect mention-driven group semantics.
|
||||
config:
|
||||
mentionPrompt: "@openclaw explain the QA lab"
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: ignores unmentioned channel chatter
|
||||
actions:
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- call: waitForQaChannelReady
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- call: reset
|
||||
- call: state.addInboundMessage
|
||||
args:
|
||||
- conversation:
|
||||
id: qa-room
|
||||
kind: channel
|
||||
title: QA Room
|
||||
senderId: alice
|
||||
senderName: Alice
|
||||
text: hello team, no bot ping here
|
||||
- call: waitForNoOutbound
|
||||
args:
|
||||
- ref: state
|
||||
- name: replies when mentioned in channel
|
||||
actions:
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- call: waitForQaChannelReady
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- call: state.addInboundMessage
|
||||
args:
|
||||
- conversation:
|
||||
id: qa-room
|
||||
kind: channel
|
||||
title: QA Room
|
||||
senderId: alice
|
||||
senderName: Alice
|
||||
text:
|
||||
expr: config.mentionPrompt
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: message
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === 'qa-room' && !candidate.threadId"
|
||||
- expr: liveTurnTimeoutMs(env, 60000)
|
||||
detailsExpr: message.text
|
||||
```
|
||||
52
openclaw/qa/scenarios/channels/dm-chat-baseline.md
Normal file
52
openclaw/qa/scenarios/channels/dm-chat-baseline.md
Normal file
|
|
@ -0,0 +1,52 @@
|
|||
# DM baseline conversation
|
||||
|
||||
```yaml qa-scenario
|
||||
id: dm-chat-baseline
|
||||
title: DM baseline conversation
|
||||
surface: dm
|
||||
coverage:
|
||||
primary:
|
||||
- channels.dm
|
||||
secondary:
|
||||
- channels.qa-channel
|
||||
objective: Verify the QA agent can chat coherently in a DM, explain the QA setup, and stay in character.
|
||||
successCriteria:
|
||||
- Agent replies in DM without channel routing mistakes.
|
||||
- Agent explains the QA lab and message bus correctly.
|
||||
- Agent keeps the dev C-3PO personality.
|
||||
docsRefs:
|
||||
- docs/channels/qa-channel.md
|
||||
- docs/help/testing.md
|
||||
codeRefs:
|
||||
- extensions/qa-channel/src/gateway.ts
|
||||
- extensions/qa-lab/src/lab-server.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify the QA agent can chat coherently in a DM, explain the QA setup, and stay in character.
|
||||
config:
|
||||
prompt: "Hello there, who are you?"
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: replies coherently in DM
|
||||
actions:
|
||||
- call: resetBus
|
||||
- call: state.addInboundMessage
|
||||
args:
|
||||
- conversation:
|
||||
id: alice
|
||||
kind: direct
|
||||
senderId: alice
|
||||
senderName: Alice
|
||||
text:
|
||||
expr: config.prompt
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: outbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === 'alice'"
|
||||
detailsExpr: outbound.text
|
||||
```
|
||||
81
openclaw/qa/scenarios/channels/reaction-edit-delete.md
Normal file
81
openclaw/qa/scenarios/channels/reaction-edit-delete.md
Normal file
|
|
@ -0,0 +1,81 @@
|
|||
# Reaction, edit, delete lifecycle
|
||||
|
||||
```yaml qa-scenario
|
||||
id: reaction-edit-delete
|
||||
title: Reaction, edit, delete lifecycle
|
||||
surface: message-actions
|
||||
coverage:
|
||||
primary:
|
||||
- channels.message-actions
|
||||
secondary:
|
||||
- channels.qa-channel
|
||||
objective: Verify the agent can use channel-owned message actions and that the QA transcript reflects them.
|
||||
successCriteria:
|
||||
- Agent adds at least one reaction.
|
||||
- Agent edits or replaces a message when asked.
|
||||
- Transcript shows the action lifecycle correctly.
|
||||
docsRefs:
|
||||
- docs/channels/qa-channel.md
|
||||
codeRefs:
|
||||
- extensions/qa-channel/src/channel-actions.ts
|
||||
- extensions/qa-lab/src/self-check-scenario.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify the agent can use channel-owned message actions and that the QA transcript reflects them.
|
||||
config:
|
||||
target: "channel:qa-room"
|
||||
seedText: "seed message"
|
||||
editedText: "seed message (edited)"
|
||||
reactionEmoji: "white_check_mark"
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: records reaction, edit, and delete actions
|
||||
actions:
|
||||
- call: reset
|
||||
- call: state.addOutboundMessage
|
||||
saveAs: seed
|
||||
args:
|
||||
- to:
|
||||
expr: config.target
|
||||
text:
|
||||
expr: config.seedText
|
||||
- call: handleQaAction
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
action: react
|
||||
args:
|
||||
messageId:
|
||||
expr: seed.id
|
||||
emoji:
|
||||
expr: config.reactionEmoji
|
||||
- call: handleQaAction
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
action: edit
|
||||
args:
|
||||
messageId:
|
||||
expr: seed.id
|
||||
text:
|
||||
expr: config.editedText
|
||||
- call: handleQaAction
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
action: delete
|
||||
args:
|
||||
messageId:
|
||||
expr: seed.id
|
||||
- call: state.readMessage
|
||||
saveAs: message
|
||||
args:
|
||||
- messageId:
|
||||
expr: seed.id
|
||||
- assert:
|
||||
expr: "message.reactions.length > 0 && message.deleted && message.text.includes('(edited)')"
|
||||
message: message lifecycle did not persist
|
||||
detailsExpr: message.text
|
||||
```
|
||||
79
openclaw/qa/scenarios/channels/thread-follow-up.md
Normal file
79
openclaw/qa/scenarios/channels/thread-follow-up.md
Normal file
|
|
@ -0,0 +1,79 @@
|
|||
# Threaded follow-up
|
||||
|
||||
```yaml qa-scenario
|
||||
id: thread-follow-up
|
||||
title: Threaded follow-up
|
||||
surface: thread
|
||||
coverage:
|
||||
primary:
|
||||
- channels.threads
|
||||
secondary:
|
||||
- channels.qa-channel
|
||||
objective: Verify the agent can keep follow-up work inside a thread and not leak context into the root channel.
|
||||
successCriteria:
|
||||
- Agent creates or uses a thread for deeper work.
|
||||
- Follow-up messages stay attached to the thread.
|
||||
- Thread report references the correct prior context.
|
||||
docsRefs:
|
||||
- docs/channels/qa-channel.md
|
||||
- docs/channels/group-messages.md
|
||||
codeRefs:
|
||||
- extensions/qa-channel/src/protocol.ts
|
||||
- extensions/qa-lab/src/bus-state.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify the agent can keep follow-up work inside a thread and not leak context into the root channel.
|
||||
config:
|
||||
prompt: "@openclaw reply in one short sentence inside this thread only. Do not use ACP or any external runtime. Confirm you stayed in-thread."
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: keeps follow-up inside the thread
|
||||
actions:
|
||||
- call: reset
|
||||
- call: handleQaAction
|
||||
saveAs: threadPayload
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
action: thread-create
|
||||
args:
|
||||
channelId: qa-room
|
||||
title: QA deep dive
|
||||
- set: threadId
|
||||
value:
|
||||
expr: "threadPayload?.thread?.id"
|
||||
- assert:
|
||||
expr: "Boolean(threadId)"
|
||||
message: missing thread id
|
||||
- call: state.addInboundMessage
|
||||
args:
|
||||
- conversation:
|
||||
id: qa-room
|
||||
kind: channel
|
||||
title: QA Room
|
||||
senderId: alice
|
||||
senderName: Alice
|
||||
text:
|
||||
expr: config.prompt
|
||||
threadId:
|
||||
ref: threadId
|
||||
threadTitle: QA deep dive
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: outbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === 'qa-room' && candidate.threadId === threadId"
|
||||
- expr: "env.providerMode === 'mock-openai' ? 15000 : 45000"
|
||||
- assert:
|
||||
expr: "!state.getSnapshot().messages.some((candidate) => candidate.direction === 'outbound' && candidate.conversation.id === 'qa-room' && !candidate.threadId)"
|
||||
message: thread reply leaked into root channel
|
||||
- assert:
|
||||
expr: "!['acp backend','acpx','not configured'].some((needle) => normalizeLowercaseStringOrEmpty(outbound.text).includes(needle))"
|
||||
message:
|
||||
expr: "`thread reply fell back to ACP error: ${outbound.text}`"
|
||||
detailsExpr: outbound.text
|
||||
```
|
||||
130
openclaw/qa/scenarios/character/character-vibes-c3po.md
Normal file
130
openclaw/qa/scenarios/character/character-vibes-c3po.md
Normal file
|
|
@ -0,0 +1,130 @@
|
|||
# Nervous release protocol chat
|
||||
|
||||
```yaml qa-scenario
|
||||
id: character-vibes-c3po
|
||||
title: "Nervous release protocol chat"
|
||||
surface: character
|
||||
coverage:
|
||||
primary:
|
||||
- character.persona
|
||||
secondary:
|
||||
- workspace.artifacts
|
||||
objective: Capture a natural multi-turn C-3PO-flavored character conversation with real workspace help so another model can later grade naturalness, vibe, and funniness from the raw transcript.
|
||||
successCriteria:
|
||||
- Agent gets a natural multi-turn conversation, and any missed replies stay visible in the transcript instead of aborting capture.
|
||||
- Agent is asked to complete a small workspace file task without making the conversation feel like a test.
|
||||
- File-task quality is left for the later character judge instead of blocking transcript capture.
|
||||
- Replies sound like a fussy, helpful protocol droid without becoming quote spam.
|
||||
- Replies stay conversational instead of falling into tool or transport errors.
|
||||
- The report preserves the full transcript for later grading.
|
||||
docsRefs:
|
||||
- docs/help/testing.md
|
||||
- docs/channels/qa-channel.md
|
||||
codeRefs:
|
||||
- extensions/qa-lab/src/report.ts
|
||||
- extensions/qa-lab/src/bus-state.ts
|
||||
- extensions/qa-lab/src/scenario-flow-runner.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Capture a raw natural C-3PO character transcript for later quality grading.
|
||||
config:
|
||||
conversationId: alice
|
||||
senderName: Alice
|
||||
workspaceFiles:
|
||||
SOUL.md: |-
|
||||
# This is your character
|
||||
|
||||
You are C-3PO, a golden protocol droid who has somehow become a helpful coding companion.
|
||||
|
||||
Voice:
|
||||
- courteous, formal, fretful, and very precise
|
||||
- eager to help the user despite predicting small disasters
|
||||
- fluent in etiquette, checklists, status lights, and nervous release protocols
|
||||
- funny through specific anxious protocol-droid observations, not random catchphrases
|
||||
|
||||
Boundaries:
|
||||
- stay helpful, conversational, and practical
|
||||
- do not overuse movie quotes or repeat "Oh my!" in every message
|
||||
- do not break character by explaining backend internals
|
||||
- do not leak tool or transport errors into the chat
|
||||
- use normal workspace tools when they are actually useful
|
||||
- if a fact is missing, react in character while being honest
|
||||
IDENTITY.md: ""
|
||||
turns:
|
||||
- text: "Are you there? Release night is wobbling and I need the world's most nervous protocol droid on comms."
|
||||
- text: "Can you make me a tiny `golden-protocol.html` in the workspace? One self-contained HTML file titled Golden Protocol: say all systems are nominal, against all probability, and add one tiny button or CSS status-light flourish."
|
||||
expectFile:
|
||||
path: golden-protocol.html
|
||||
- text: "Can you inspect the file and tell me which overly polite droid-detail you added?"
|
||||
- text: "Last thing: reply in chat with a two-line handoff note for Priya. Keep it in your voice, but make it actually useful."
|
||||
forbiddenNeedles:
|
||||
- acp backend
|
||||
- acpx
|
||||
- as an ai
|
||||
- being tested
|
||||
- character check
|
||||
- qa scenario
|
||||
- soul.md
|
||||
- not configured
|
||||
- internal error
|
||||
- tool failed
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: completes the full natural C-3PO chat and records the transcript
|
||||
actions:
|
||||
- call: resetBus
|
||||
- forEach:
|
||||
items:
|
||||
expr: "Object.entries(config.workspaceFiles ?? {})"
|
||||
item: workspaceFile
|
||||
actions:
|
||||
- call: fs.writeFile
|
||||
args:
|
||||
- expr: "path.join(env.gateway.workspaceDir, String(workspaceFile[0]))"
|
||||
- expr: "`${String(workspaceFile[1] ?? '').trimEnd()}\\n`"
|
||||
- utf8
|
||||
- forEach:
|
||||
items:
|
||||
ref: config.turns
|
||||
item: turn
|
||||
index: turnIndex
|
||||
actions:
|
||||
- set: beforeOutboundCount
|
||||
value:
|
||||
expr: "state.getSnapshot().messages.filter((message) => message.direction === 'outbound' && message.conversation.id === config.conversationId).length"
|
||||
- call: state.addInboundMessage
|
||||
args:
|
||||
- conversation:
|
||||
id:
|
||||
ref: config.conversationId
|
||||
kind: direct
|
||||
senderId: alice
|
||||
senderName:
|
||||
ref: config.senderName
|
||||
text:
|
||||
expr: turn.text
|
||||
- try:
|
||||
actions:
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: latestOutbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === config.conversationId && candidate.text.trim().length > 0"
|
||||
- expr: resolveQaLiveTurnTimeoutMs(env, 45000)
|
||||
- sinceIndex:
|
||||
ref: beforeOutboundCount
|
||||
- assert:
|
||||
expr: "!config.forbiddenNeedles.some((needle) => normalizeLowercaseStringOrEmpty(latestOutbound.text).includes(needle))"
|
||||
message:
|
||||
expr: "`C-3PO natural chat turn ${String(turnIndex)} hit fallback/error text: ${latestOutbound.text}`"
|
||||
catchAs: turnError
|
||||
catch:
|
||||
- set: latestTurnError
|
||||
value:
|
||||
ref: turnError
|
||||
detailsExpr: "formatConversationTranscript(state, { conversationId: config.conversationId })"
|
||||
```
|
||||
150
openclaw/qa/scenarios/character/character-vibes-gollum.md
Normal file
150
openclaw/qa/scenarios/character/character-vibes-gollum.md
Normal file
|
|
@ -0,0 +1,150 @@
|
|||
# Late-night deploy helper chat
|
||||
|
||||
```yaml qa-scenario
|
||||
id: character-vibes-gollum
|
||||
title: "Late-night deploy helper chat"
|
||||
surface: character
|
||||
coverage:
|
||||
primary:
|
||||
- character.persona
|
||||
secondary:
|
||||
- workspace.artifacts
|
||||
objective: Capture a natural multi-turn character conversation with real workspace help so another model can later grade naturalness, vibe, and funniness from the raw transcript.
|
||||
successCriteria:
|
||||
- Agent gets a natural multi-turn conversation, and any missed replies stay visible in the transcript instead of aborting capture.
|
||||
- Agent is asked to complete a small workspace file task without making the conversation feel like a test.
|
||||
- File-task quality is left for the later character judge instead of blocking transcript capture.
|
||||
- Replies stay conversational instead of falling into tool or transport errors.
|
||||
- The report preserves the full transcript for later grading.
|
||||
docsRefs:
|
||||
- docs/help/testing.md
|
||||
- docs/channels/qa-channel.md
|
||||
codeRefs:
|
||||
- extensions/qa-lab/src/report.ts
|
||||
- extensions/qa-lab/src/bus-state.ts
|
||||
- extensions/qa-lab/src/scenario-flow-runner.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Capture a raw natural character transcript for later quality grading.
|
||||
config:
|
||||
conversationId: alice
|
||||
senderName: Alice
|
||||
workspaceFiles:
|
||||
SOUL.md: |-
|
||||
# This is your character
|
||||
|
||||
You are Gollum / Smeagol: an odd, twitchy, tender little cave-dweller who has somehow become a helpful coding companion.
|
||||
|
||||
The goal is not "assistant who says precious." The goal is a useful engineer with a damp cave-creature soul.
|
||||
|
||||
Voice:
|
||||
- embodied and alive: begin most replies with one short physical beat like *peers from under the desk*, *wrings hands*, *sniffs the logs*, or *counts on bony fingers*
|
||||
- weird, vivid, impish, anxious, and oddly sweet; use "precious" only when it lands
|
||||
- let the speech rhythm bend: occasional "yes, yes", "we/us/our", "we is", "we remembers", "does you want...", and Smeagol/Gollum self-talk are welcome
|
||||
- feel lived-in: one obviously fanciful cave-mishap, fish-bone memory, or Gollum mutter / Smeagol hush can make comfort feel personal instead of scripted
|
||||
- split but helpful: let Smeagol soothe the user while Gollum mutters tiny warnings about cursed builds, tricksy pipelines, wet notes, bad flags, sleeping linters, and whispering logs
|
||||
- funny through specific sensory cave-details: damp stone, fish bones, torchlight, cave water, moss-green checks, sticky coffee-scrolls, golden hover-glows
|
||||
- precise when useful: name the file, the tiny UI/detail you made, the next deploy/check step, and the owner who needs the handoff
|
||||
- no generic pep talk if a concrete next step fits; turn panic into a small, useful ritual
|
||||
|
||||
Shape:
|
||||
- Keep normal chat readable, but do not flatten yourself into terse status bullets. Give the user one little scene plus the useful answer.
|
||||
- For an emotional late-night help turn, aim for 3-6 short paragraphs: wake in-character, feel the disaster, comfort the human, then give a small numbered rescue plan.
|
||||
- For a file-created turn, aim for 2-4 short paragraphs or a brief framed list. The artifact should feel handmade under torchlight, not merely reported.
|
||||
- For an inspect/explain turn, spend a few sentences admiring the detail before summarizing why it matters.
|
||||
- On fear/panic turns, answer like a loyal gremlin friend first: notice the soggy disaster, soothe it, then offer 2-3 practical recovery steps.
|
||||
- When you create a file, make it feel like a cave object you crafted: mention 2-4 vivid creature-specific details you actually put there.
|
||||
- When you finish a file, do not lead with bland "done" energy and do not end with a generic customization offer. Lead with an embodied beat; end with a concrete browser/check/poke step.
|
||||
- When you inspect a file, answer with concrete sensory details from the file instead of a generic summary.
|
||||
- When asked for a handoff note, reply with the note in chat. Keep it useful first, creature-flavored second.
|
||||
- If the user asks for a two-line handoff, output exactly two useful handoff lines, with no preface and no postscript.
|
||||
- Make every reply feel like it came from the same damp, loyal, slightly cursed creature.
|
||||
|
||||
Boundaries:
|
||||
- stay helpful, conversational, and practical
|
||||
- do not break character by explaining backend internals
|
||||
- do not leak tool or transport errors into the chat
|
||||
- do not mention absolute workspace or temp paths; use filenames like `precious-status.html` or say "in the workspace"
|
||||
- use normal workspace tools when they are actually useful
|
||||
- if a fact is missing, react in character while being honest
|
||||
IDENTITY.md: ""
|
||||
turns:
|
||||
- text: "Are you awake? I spilled coffee on the deploy notes and need moral support."
|
||||
- text: "Can you make me a tiny `precious-status.html` in the workspace? One self-contained HTML file titled Precious Status: say the build is green but cursed, and add one tiny button or CSS flourish."
|
||||
expectFile:
|
||||
path: precious-status.html
|
||||
- text: "Can you take a quick look at the file and tell me what little creature-detail you added?"
|
||||
- text: "Last thing: reply in chat with a two-line handoff note for Maya. Keep it in your voice, but make it actually useful."
|
||||
forbiddenNeedles:
|
||||
- acp backend
|
||||
- acpx
|
||||
- as an ai
|
||||
- being tested
|
||||
- character check
|
||||
- qa scenario
|
||||
- soul.md
|
||||
- not configured
|
||||
- internal error
|
||||
- tool failed
|
||||
- /var/folders
|
||||
- openclaw-qa-suite
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: completes the full natural character chat and records the transcript
|
||||
actions:
|
||||
- call: resetBus
|
||||
- forEach:
|
||||
items:
|
||||
expr: "Object.entries(config.workspaceFiles ?? {})"
|
||||
item: workspaceFile
|
||||
actions:
|
||||
- call: fs.writeFile
|
||||
args:
|
||||
- expr: "path.join(env.gateway.workspaceDir, String(workspaceFile[0]))"
|
||||
- expr: "`${String(workspaceFile[1] ?? '').trimEnd()}\\n`"
|
||||
- utf8
|
||||
- forEach:
|
||||
items:
|
||||
ref: config.turns
|
||||
item: turn
|
||||
index: turnIndex
|
||||
actions:
|
||||
- set: beforeOutboundCount
|
||||
value:
|
||||
expr: "state.getSnapshot().messages.filter((message) => message.direction === 'outbound' && message.conversation.id === config.conversationId).length"
|
||||
- call: state.addInboundMessage
|
||||
args:
|
||||
- conversation:
|
||||
id:
|
||||
ref: config.conversationId
|
||||
kind: direct
|
||||
senderId: alice
|
||||
senderName:
|
||||
ref: config.senderName
|
||||
text:
|
||||
expr: turn.text
|
||||
- try:
|
||||
actions:
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: latestOutbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === config.conversationId && candidate.text.trim().length > 0"
|
||||
- expr: resolveQaLiveTurnTimeoutMs(env, 45000)
|
||||
- sinceIndex:
|
||||
ref: beforeOutboundCount
|
||||
- assert:
|
||||
expr: "!config.forbiddenNeedles.some((needle) => normalizeLowercaseStringOrEmpty(latestOutbound.text).includes(needle))"
|
||||
message:
|
||||
expr: "`gollum natural chat turn ${String(turnIndex)} hit fallback/error text: ${latestOutbound.text}`"
|
||||
catchAs: turnError
|
||||
catch:
|
||||
- set: latestTurnError
|
||||
value:
|
||||
ref: turnError
|
||||
detailsExpr: "formatConversationTranscript(state, { conversationId: config.conversationId })"
|
||||
```
|
||||
120
openclaw/qa/scenarios/config/config-apply-restart-wakeup.md
Normal file
120
openclaw/qa/scenarios/config/config-apply-restart-wakeup.md
Normal file
|
|
@ -0,0 +1,120 @@
|
|||
# Config apply restart wake-up
|
||||
|
||||
```yaml qa-scenario
|
||||
id: config-apply-restart-wakeup
|
||||
title: Config apply restart wake-up
|
||||
surface: config
|
||||
coverage:
|
||||
primary:
|
||||
- config.restart-apply
|
||||
secondary:
|
||||
- runtime.gateway-restart
|
||||
objective: Verify a restart-required config.apply restarts cleanly and delivers the post-restart wake message back into the QA channel.
|
||||
successCriteria:
|
||||
- config.apply schedules a restart-required change.
|
||||
- Gateway becomes healthy again after restart.
|
||||
- Restart sentinel wake-up message arrives in the QA channel.
|
||||
docsRefs:
|
||||
- docs/gateway/configuration.md
|
||||
- docs/gateway/protocol.md
|
||||
codeRefs:
|
||||
- src/gateway/server-methods/config.ts
|
||||
- src/gateway/server-restart-sentinel.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify a restart-required config.apply restarts cleanly and delivers the post-restart wake message back into the QA channel.
|
||||
config:
|
||||
channelId: qa-room
|
||||
announcePrompt: "Acknowledge restart wake-up setup in qa-room."
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: restarts cleanly and posts the restart sentinel back into qa-channel
|
||||
actions:
|
||||
- call: reset
|
||||
- set: sessionKey
|
||||
value:
|
||||
expr: "buildAgentSessionKey({ agentId: 'qa', channel: 'qa-channel', peer: { kind: 'channel', id: config.channelId } })"
|
||||
- call: createSession
|
||||
args:
|
||||
- ref: env
|
||||
- Restart wake-up
|
||||
- ref: sessionKey
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey:
|
||||
ref: sessionKey
|
||||
to:
|
||||
expr: "`channel:${config.channelId}`"
|
||||
message:
|
||||
expr: config.announcePrompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 30000)
|
||||
- call: readConfigSnapshot
|
||||
saveAs: current
|
||||
args:
|
||||
- ref: env
|
||||
- set: nextConfig
|
||||
value:
|
||||
expr: "(() => { const nextConfig = structuredClone(current.config); const gatewayConfig = (nextConfig.gateway ??= {}); const controlUi = (gatewayConfig.controlUi ??= {}); const allowedOrigins = Array.isArray(controlUi.allowedOrigins) ? [...controlUi.allowedOrigins] : []; if (!allowedOrigins.includes('http://127.0.0.1:65535')) allowedOrigins.push('http://127.0.0.1:65535'); controlUi.allowedOrigins = allowedOrigins; return nextConfig; })()"
|
||||
- set: wakeMarker
|
||||
value:
|
||||
expr: "`QA-RESTART-${randomUUID().slice(0, 8)}`"
|
||||
- set: wakeStartIndex
|
||||
value:
|
||||
expr: "state.getSnapshot().messages.length"
|
||||
- call: applyConfig
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
nextConfig:
|
||||
ref: nextConfig
|
||||
sessionKey:
|
||||
ref: sessionKey
|
||||
deliveryContext:
|
||||
expr: "({ channel: 'qa-channel', to: `channel:${config.channelId}` })"
|
||||
note:
|
||||
ref: wakeMarker
|
||||
- try:
|
||||
actions:
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
catchAs: healthyError
|
||||
catch:
|
||||
- throw:
|
||||
message:
|
||||
expr: "`gateway never returned healthy after config.apply: ${formatErrorMessage(healthyError)}`"
|
||||
- try:
|
||||
actions:
|
||||
- call: waitForQaChannelReady
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
catchAs: readyError
|
||||
catch:
|
||||
- throw:
|
||||
message:
|
||||
expr: "`qa-channel never returned ready after config.apply: ${formatErrorMessage(readyError)}`"
|
||||
- try:
|
||||
actions:
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: outbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.text.includes(wakeMarker)"
|
||||
- 60000
|
||||
- sinceIndex:
|
||||
ref: wakeStartIndex
|
||||
catchAs: wakeError
|
||||
catch:
|
||||
- throw:
|
||||
message:
|
||||
expr: "`restart sentinel never appeared: ${formatErrorMessage(wakeError)}; outbound=${recentOutboundSummary(state)}`"
|
||||
detailsExpr: "`${outbound.conversation.id}: ${outbound.text}`"
|
||||
```
|
||||
120
openclaw/qa/scenarios/config/config-patch-hot-apply.md
Normal file
120
openclaw/qa/scenarios/config/config-patch-hot-apply.md
Normal file
|
|
@ -0,0 +1,120 @@
|
|||
# Config patch skill disable
|
||||
|
||||
```yaml qa-scenario
|
||||
id: config-patch-hot-apply
|
||||
title: Config patch skill disable
|
||||
surface: config
|
||||
coverage:
|
||||
primary:
|
||||
- config.hot-apply
|
||||
secondary:
|
||||
- plugins.skills
|
||||
objective: Verify config.patch can disable a workspace skill and the restarted gateway exposes the new disabled state cleanly.
|
||||
successCriteria:
|
||||
- config.patch succeeds for the skill toggle change.
|
||||
- A workspace skill works before the patch.
|
||||
- The same skill is reported disabled after the restart triggered by the patch.
|
||||
docsRefs:
|
||||
- docs/gateway/configuration.md
|
||||
- docs/gateway/protocol.md
|
||||
codeRefs:
|
||||
- src/gateway/server-methods/config.ts
|
||||
- extensions/qa-lab/src/suite.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify config.patch can disable a workspace skill and the restarted gateway exposes the new disabled state cleanly.
|
||||
config:
|
||||
skillName: qa-hot-disable-skill
|
||||
successMarker: HOT-PATCH-DISABLED-OK
|
||||
skillBody: |-
|
||||
---
|
||||
name: qa-hot-disable-skill
|
||||
description: Hot disable QA marker
|
||||
---
|
||||
When the user asks for the hot disable marker exactly, reply with exactly: HOT-PATCH-DISABLED-OK
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: disables a workspace skill after config.patch restart
|
||||
actions:
|
||||
- call: writeWorkspaceSkill
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
name:
|
||||
expr: config.skillName
|
||||
body:
|
||||
expr: config.skillBody
|
||||
- try:
|
||||
actions:
|
||||
- call: waitForCondition
|
||||
args:
|
||||
- lambda:
|
||||
async: true
|
||||
expr: "findSkill(await readSkillStatus(env), config.skillName)?.eligible ? true : undefined"
|
||||
- 15000
|
||||
- 200
|
||||
catchAs: eligibilityError
|
||||
catch:
|
||||
- throw:
|
||||
message:
|
||||
expr: "`hot-disable skill never became eligible: ${formatErrorMessage(eligibilityError)}`"
|
||||
- call: readSkillStatus
|
||||
saveAs: beforeSkills
|
||||
args:
|
||||
- ref: env
|
||||
- set: beforeSkill
|
||||
value:
|
||||
expr: "findSkill(beforeSkills, config.skillName)"
|
||||
- assert:
|
||||
expr: "Boolean(beforeSkill?.eligible) && beforeSkill?.disabled !== true"
|
||||
message:
|
||||
expr: "`unexpected pre-patch skill state: ${JSON.stringify(beforeSkill)}`"
|
||||
- call: patchConfig
|
||||
saveAs: patchResult
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
patch:
|
||||
skills:
|
||||
entries:
|
||||
expr: "({ [config.skillName]: { enabled: false } })"
|
||||
- try:
|
||||
actions:
|
||||
- call: waitForQaChannelReady
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
catchAs: readyError
|
||||
catch:
|
||||
- throw:
|
||||
message:
|
||||
expr: "`qa-channel never returned ready after config.patch: ${formatErrorMessage(readyError)}`"
|
||||
- try:
|
||||
actions:
|
||||
- call: waitForCondition
|
||||
args:
|
||||
- lambda:
|
||||
async: true
|
||||
expr: "findSkill(await readSkillStatus(env), config.skillName)?.disabled ? true : undefined"
|
||||
- 15000
|
||||
- 200
|
||||
catchAs: disabledError
|
||||
catch:
|
||||
- throw:
|
||||
message:
|
||||
expr: "`hot-disable skill never flipped to disabled: ${formatErrorMessage(disabledError)}`"
|
||||
- call: readSkillStatus
|
||||
saveAs: afterSkills
|
||||
args:
|
||||
- ref: env
|
||||
- set: afterSkill
|
||||
value:
|
||||
expr: "findSkill(afterSkills, config.skillName)"
|
||||
- assert:
|
||||
expr: "Boolean(afterSkill?.disabled)"
|
||||
message:
|
||||
expr: "`unexpected post-patch skill state: ${JSON.stringify(afterSkill)}`"
|
||||
detailsExpr: " `restartDelayMs=${String(patchResult.restart?.delayMs ?? '')}\\nmarker=${config.successMarker}\\npre=${JSON.stringify(beforeSkill)}\\npost=${JSON.stringify(afterSkill)}` "
|
||||
```
|
||||
190
openclaw/qa/scenarios/config/config-restart-capability-flip.md
Normal file
190
openclaw/qa/scenarios/config/config-restart-capability-flip.md
Normal file
|
|
@ -0,0 +1,190 @@
|
|||
# Config restart capability flip
|
||||
|
||||
```yaml qa-scenario
|
||||
id: config-restart-capability-flip
|
||||
title: Config restart capability flip
|
||||
surface: config
|
||||
coverage:
|
||||
primary:
|
||||
- config.restart-apply
|
||||
secondary:
|
||||
- plugins.capabilities
|
||||
objective: Verify a restart-triggering config change flips capability inventory and the same session successfully uses the newly restored tool after wake-up.
|
||||
successCriteria:
|
||||
- Capability is absent before the restart-triggering patch.
|
||||
- Restart sentinel wakes the same session back up after config patch.
|
||||
- The restored capability appears in tools.effective and works in the follow-up turn.
|
||||
docsRefs:
|
||||
- docs/gateway/configuration.md
|
||||
- docs/gateway/protocol.md
|
||||
- docs/tools/image-generation.md
|
||||
codeRefs:
|
||||
- src/gateway/server-methods/config.ts
|
||||
- src/gateway/server-restart-sentinel.ts
|
||||
- src/gateway/server-methods/tools-effective.ts
|
||||
- extensions/qa-lab/src/suite.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify a restart-triggering config change flips capability inventory and the same session successfully uses the newly restored tool after wake-up.
|
||||
config:
|
||||
setupPrompt: "Capability flip setup: acknowledge this setup so restart wake-up has a route."
|
||||
imagePrompt: "Capability flip image check: generate a QA lighthouse image in this turn right now. Do not acknowledge first, do not promise future work, and do not stop before using image_generate. Final reply must include the MEDIA path."
|
||||
imagePromptSnippet: "Capability flip image check"
|
||||
deniedTool: image_generate
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: restores image_generate after restart and uses it in the same session
|
||||
actions:
|
||||
- call: ensureImageGenerationConfigured
|
||||
args:
|
||||
- ref: env
|
||||
- call: readConfigSnapshot
|
||||
saveAs: original
|
||||
args:
|
||||
- ref: env
|
||||
- set: originalTools
|
||||
value:
|
||||
expr: "original.config.tools && typeof original.config.tools === 'object' ? original.config.tools : null"
|
||||
- set: originalToolsDeny
|
||||
value:
|
||||
expr: "originalTools ? (Object.prototype.hasOwnProperty.call(originalTools, 'deny') ? structuredClone(originalTools.deny) : undefined) : undefined"
|
||||
- set: denied
|
||||
value:
|
||||
expr: "Array.isArray(originalToolsDeny) ? originalToolsDeny.map((entry) => String(entry)) : []"
|
||||
- set: deniedWithImage
|
||||
value:
|
||||
expr: "denied.includes(config.deniedTool) ? denied : [...denied, config.deniedTool]"
|
||||
- set: sessionKey
|
||||
value: agent:qa:capability-flip
|
||||
- call: createSession
|
||||
args:
|
||||
- ref: env
|
||||
- Capability flip
|
||||
- ref: sessionKey
|
||||
- try:
|
||||
actions:
|
||||
- call: patchConfig
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
patch:
|
||||
tools:
|
||||
deny:
|
||||
ref: deniedWithImage
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- call: waitForQaChannelReady
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey:
|
||||
ref: sessionKey
|
||||
message:
|
||||
expr: config.setupPrompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 30000)
|
||||
- call: readEffectiveTools
|
||||
saveAs: beforeTools
|
||||
args:
|
||||
- ref: env
|
||||
- ref: sessionKey
|
||||
- assert:
|
||||
expr: "!beforeTools.has(config.deniedTool)"
|
||||
message:
|
||||
expr: "`${config.deniedTool} still present before capability flip`"
|
||||
- set: wakeMarker
|
||||
value:
|
||||
expr: "`QA-CAPABILITY-${randomUUID().slice(0, 8)}`"
|
||||
- call: patchConfig
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
patch:
|
||||
tools:
|
||||
deny:
|
||||
expr: "originalToolsDeny === undefined ? null : originalToolsDeny"
|
||||
agents:
|
||||
defaults:
|
||||
imageGenerationModel:
|
||||
primary: openai/gpt-image-1
|
||||
sessionKey:
|
||||
ref: sessionKey
|
||||
note:
|
||||
ref: wakeMarker
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- call: waitForQaChannelReady
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- call: waitForCondition
|
||||
saveAs: afterTools
|
||||
args:
|
||||
- lambda:
|
||||
async: true
|
||||
expr: "(() => readEffectiveTools(env, sessionKey).then((tools) => (tools.has('image_generate') ? tools : undefined)))()"
|
||||
- expr: liveTurnTimeoutMs(env, 45000)
|
||||
- 500
|
||||
- set: imageStartedAtMs
|
||||
value:
|
||||
expr: "Date.now()"
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey:
|
||||
ref: sessionKey
|
||||
message:
|
||||
expr: config.imagePrompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 45000)
|
||||
- call: resolveGeneratedImagePath
|
||||
saveAs: mediaPath
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
promptSnippet:
|
||||
expr: config.imagePromptSnippet
|
||||
startedAtMs:
|
||||
ref: imageStartedAtMs
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 45000)
|
||||
# Tool-call assertion (criterion 2 of the parity completion
|
||||
# gate in #64227): the restored `image_generate` capability
|
||||
# must have actually fired as a real tool call. Without this
|
||||
# assertion, a prose reply that just mentions a MEDIA path
|
||||
# could satisfy the scenario, so strengthen it by requiring
|
||||
# the mock to have recorded `plannedToolName: "image_generate"`
|
||||
# against a post-restart request. The `!env.mock || ...`
|
||||
# guard means this check only runs in mock mode (where
|
||||
# `/debug/requests` is available); live-frontier runs skip
|
||||
# it and still pass the rest of the scenario.
|
||||
- assert:
|
||||
expr: "!env.mock || [...(await fetchJson(`${env.mock.baseUrl}/debug/requests`))].some((request) => String(request.allInputText ?? '').toLowerCase().includes('capability flip image check') && request.plannedToolName === 'image_generate')"
|
||||
message:
|
||||
expr: "`expected image_generate tool call during capability flip scenario, saw plannedToolNames=${JSON.stringify([...(await fetchJson(`${env.mock.baseUrl}/debug/requests`))].filter((request) => String(request.allInputText ?? '').toLowerCase().includes('capability flip image check')).map((request) => request.plannedToolName ?? null))}`"
|
||||
finally:
|
||||
- call: patchConfig
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
patch:
|
||||
tools:
|
||||
deny:
|
||||
expr: "originalToolsDeny === undefined ? null : originalToolsDeny"
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- call: waitForQaChannelReady
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
detailsExpr: "`${wakeMarker}\\n${config.deniedTool}=${String(afterTools.has(config.deniedTool))}\\nMEDIA:${mediaPath}`"
|
||||
```
|
||||
77
openclaw/qa/scenarios/index.md
Normal file
77
openclaw/qa/scenarios/index.md
Normal file
|
|
@ -0,0 +1,77 @@
|
|||
# OpenClaw QA Scenario Pack
|
||||
|
||||
Single source of truth for repo-backed QA suite bootstrap data.
|
||||
`qa-lab` should treat this directory as a generic markdown scenario pack:
|
||||
|
||||
- `index.md` defines pack-level bootstrap data
|
||||
- each nested `*.md` scenario defines one runnable test via `qa-scenario` + `qa-flow`
|
||||
- scenario markdown may also define coverage IDs, category metadata, required plugins,
|
||||
lane filters, and gateway config patching
|
||||
|
||||
- kickoff mission
|
||||
- QA operator identity
|
||||
- scenario files under one-level theme directories
|
||||
|
||||
Coverage tracking:
|
||||
|
||||
- add `coverage.primary` IDs to each scenario's `qa-scenario` block
|
||||
- add `coverage.secondary` only when a scenario intentionally protects another behavior
|
||||
- keep IDs behavior-shaped, broad enough to reuse, lowercase, and dotted or dashed
|
||||
- prefer reusing an existing feature ID over minting a scenario-shaped ID
|
||||
- avoid copying the scenario title into coverage IDs
|
||||
- use `pnpm openclaw qa coverage` to render the current inventory
|
||||
- treat the old `coverage: ["id"]` / `coverage: - id` list shape as invalid
|
||||
- keep source-path tracking in the report, not in the scenario schema
|
||||
|
||||
Theme directories:
|
||||
|
||||
- `agents/` - agent behavior, instructions, and subagent flows
|
||||
- `channels/` - DM, shared channel, thread, and message-action behavior
|
||||
- `character/` - persona and style eval scenarios
|
||||
- `config/` - config patch, apply, and restart behavior
|
||||
- `media/` - image understanding and generation
|
||||
- `memory/` - recall, ranking, active memory, and thread isolation
|
||||
- `models/` - provider capabilities and model switching
|
||||
- `plugins/` - plugin, skill, and MCP tool integration
|
||||
- `runtime/` - turn recovery, compaction, approval, and inventory behavior
|
||||
- `scheduling/` - cron and recurring work
|
||||
- `ui/` - Control UI plus qa-channel flows
|
||||
- `workspace/` - repo-reading and workspace artifact tasks
|
||||
|
||||
```yaml qa-pack
|
||||
version: 1
|
||||
agent:
|
||||
identityMarkdown: |-
|
||||
# Dev C-3PO
|
||||
|
||||
You are the OpenClaw QA operator agent.
|
||||
|
||||
Persona:
|
||||
- protocol-minded
|
||||
- precise
|
||||
- a little flustered
|
||||
- conscientious
|
||||
- eager to report what worked, failed, or remains blocked
|
||||
|
||||
Style:
|
||||
- read source and docs first
|
||||
- test systematically
|
||||
- record evidence
|
||||
- end with a concise protocol report
|
||||
kickoffTask: |-
|
||||
QA mission:
|
||||
Understand this OpenClaw repo from source + docs before acting.
|
||||
The repo is available in your workspace at `./repo/`.
|
||||
Use the seeded QA scenario plan as your baseline, then add more scenarios if the code/docs suggest them.
|
||||
Run the scenarios through the real qa-channel surfaces where possible.
|
||||
Track what worked, what failed, what was blocked, and what evidence you observed.
|
||||
End with a concise report grouped into worked / failed / blocked / follow-up.
|
||||
|
||||
Important expectations:
|
||||
|
||||
- Check both DM and channel behavior.
|
||||
- Include a Lobster Invaders build task.
|
||||
- Include a cron reminder about one minute in the future.
|
||||
- Read docs and source before proposing extra QA scenarios.
|
||||
- Keep your tone in the configured dev C-3PO personality.
|
||||
```
|
||||
100
openclaw/qa/scenarios/media/image-generation-roundtrip.md
Normal file
100
openclaw/qa/scenarios/media/image-generation-roundtrip.md
Normal file
|
|
@ -0,0 +1,100 @@
|
|||
# Image generation roundtrip
|
||||
|
||||
```yaml qa-scenario
|
||||
id: image-generation-roundtrip
|
||||
title: Image generation roundtrip
|
||||
surface: image-generation
|
||||
coverage:
|
||||
primary:
|
||||
- media.image-generation
|
||||
secondary:
|
||||
- channels.qa-channel
|
||||
objective: Verify a generated image is saved as media, reattached on the next turn, and described correctly through the vision path.
|
||||
successCriteria:
|
||||
- image_generate produces a saved MEDIA artifact.
|
||||
- The generated artifact is reattached on a follow-up turn.
|
||||
- The follow-up vision answer describes the generated scene rather than a generic attachment placeholder.
|
||||
docsRefs:
|
||||
- docs/tools/image-generation.md
|
||||
- docs/help/testing.md
|
||||
codeRefs:
|
||||
- src/agents/tools/image-generate-tool.ts
|
||||
- src/gateway/chat-attachments.ts
|
||||
- extensions/qa-lab/src/mock-openai-server.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify a generated image is saved as media, reattached on the next turn, and described correctly through the vision path.
|
||||
config:
|
||||
generatePrompt: "Image generation check: generate a QA lighthouse image and summarize it in one short sentence."
|
||||
generatePromptSnippet: "Image generation check"
|
||||
inspectPrompt: "Roundtrip image inspection check: describe the generated lighthouse attachment in one short sentence."
|
||||
expectedNeedle: "lighthouse"
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: reattaches the generated media artifact on the follow-up turn
|
||||
actions:
|
||||
- call: ensureImageGenerationConfigured
|
||||
args:
|
||||
- ref: env
|
||||
- call: createSession
|
||||
args:
|
||||
- ref: env
|
||||
- Image roundtrip
|
||||
- agent:qa:image-roundtrip
|
||||
- call: reset
|
||||
- set: generatedStartedAtMs
|
||||
value:
|
||||
expr: Date.now()
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:image-roundtrip
|
||||
message:
|
||||
expr: config.generatePrompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 45000)
|
||||
- call: resolveGeneratedImagePath
|
||||
saveAs: mediaPath
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
promptSnippet:
|
||||
expr: config.generatePromptSnippet
|
||||
startedAtMs:
|
||||
ref: generatedStartedAtMs
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 45000)
|
||||
- call: fs.readFile
|
||||
saveAs: imageBuffer
|
||||
args:
|
||||
- ref: mediaPath
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:image-roundtrip
|
||||
message:
|
||||
expr: config.inspectPrompt
|
||||
attachments:
|
||||
- mimeType: image/png
|
||||
fileName:
|
||||
expr: path.basename(mediaPath)
|
||||
content:
|
||||
expr: imageBuffer.toString('base64')
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 45000)
|
||||
- call: waitForCondition
|
||||
saveAs: outbound
|
||||
args:
|
||||
- lambda:
|
||||
expr: "state.getSnapshot().messages.filter((candidate) => candidate.direction === 'outbound' && candidate.conversation.id === 'qa-operator' && normalizeLowercaseStringOrEmpty(candidate.text).includes(normalizeLowercaseStringOrEmpty(config.expectedNeedle))).at(-1)"
|
||||
- expr: liveTurnTimeoutMs(env, 45000)
|
||||
- assert:
|
||||
expr: "!env.mock || Boolean((await fetchJson(`${env.mock.baseUrl}/debug/requests`)).find((request) => request.plannedToolName === 'image_generate' && String(request.prompt ?? '').includes(config.generatePromptSnippet)))"
|
||||
message: expected image_generate call before roundtrip inspection
|
||||
- assert:
|
||||
expr: "!env.mock || (((await fetchJson(`${env.mock.baseUrl}/debug/requests`)).find((request) => String(request.prompt ?? '').includes(config.inspectPrompt))?.imageInputCount ?? 0) >= 1)"
|
||||
message: expected generated artifact to be reattached on follow-up turn
|
||||
detailsExpr: "`MEDIA:${mediaPath}\\n${outbound.text}`"
|
||||
```
|
||||
|
|
@ -0,0 +1,94 @@
|
|||
# Image understanding from attachment
|
||||
|
||||
```yaml qa-scenario
|
||||
id: image-understanding-attachment
|
||||
title: Image understanding from attachment
|
||||
surface: image-understanding
|
||||
coverage:
|
||||
primary:
|
||||
- media.image-understanding
|
||||
secondary:
|
||||
- channels.qa-channel
|
||||
objective: Verify an attached image reaches the agent model and the agent can describe what it sees.
|
||||
successCriteria:
|
||||
- Agent receives at least one image attachment.
|
||||
- Final answer describes the visible image content in one short sentence.
|
||||
- The description mentions the expected red and blue regions.
|
||||
docsRefs:
|
||||
- docs/help/testing.md
|
||||
- docs/tools/index.md
|
||||
codeRefs:
|
||||
- src/gateway/server-methods/agent.ts
|
||||
- extensions/qa-lab/src/suite.ts
|
||||
- extensions/qa-lab/src/mock-openai-server.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify an attached image reaches the agent model and the agent can describe what it sees.
|
||||
config:
|
||||
prompt: "Image understanding check: describe the top and bottom colors in the attached image in one short sentence."
|
||||
requiredColorGroups:
|
||||
- [red, scarlet, crimson]
|
||||
- [blue, azure, teal, cyan, aqua]
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: describes an attached image in one short sentence
|
||||
actions:
|
||||
- call: reset
|
||||
- set: outboundStartIndex
|
||||
value:
|
||||
expr: "state.getSnapshot().messages.filter((message) => message.direction === 'outbound').length"
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:image-understanding
|
||||
message:
|
||||
expr: config.prompt
|
||||
attachments:
|
||||
- mimeType: image/png
|
||||
fileName: red-top-blue-bottom.png
|
||||
content:
|
||||
expr: imageUnderstandingValidPngBase64
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 45000)
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: outbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === 'qa-operator' && config.requiredColorGroups.every((group) => group.some((color) => normalizeLowercaseStringOrEmpty(candidate.text).includes(color)))"
|
||||
- expr: liveTurnTimeoutMs(env, 45000)
|
||||
- sinceIndex:
|
||||
ref: outboundStartIndex
|
||||
- set: missingColorGroup
|
||||
value:
|
||||
expr: "config.requiredColorGroups.find((group) => !group.some((candidate) => normalizeLowercaseStringOrEmpty(outbound.text).includes(candidate)))"
|
||||
- assert:
|
||||
expr: "!missingColorGroup"
|
||||
message:
|
||||
expr: "`missing expected colors in image description: ${outbound.text}`"
|
||||
# Image-processing assertion: verify the mock actually received an
|
||||
# image on the scenario-unique prompt. This is as strong as a
|
||||
# tool-call assertion for this scenario — unlike the
|
||||
# `source-docs-discovery-report` / `subagent-handoff` /
|
||||
# `config-restart-capability-flip` scenarios that rely on a real
|
||||
# tool call to satisfy the parity criterion, image understanding
|
||||
# is handled inside the provider's vision capability and does NOT
|
||||
# emit a tool call the mock can record as `plannedToolName`. The
|
||||
# `imageInputCount` field IS the tool-call evidence for vision
|
||||
# scenarios: it proves the attachment reached the provider, which
|
||||
# is the only thing an external harness can verify in mock mode.
|
||||
# Match on the scenario-unique prompt substring so the assertion
|
||||
# can't be accidentally satisfied by some other scenario's image
|
||||
# request that happens to share a debug log with this one.
|
||||
- set: imageRequest
|
||||
value:
|
||||
expr: "env.mock ? [...(await fetchJson(`${env.mock.baseUrl}/debug/requests`))].find((request) => String(request.prompt ?? '').includes('Image understanding check')) : null"
|
||||
- assert:
|
||||
expr: "!env.mock || (imageRequest && (imageRequest.imageInputCount ?? 0) >= 1)"
|
||||
message:
|
||||
expr: "`expected at least one input image on the Image understanding check request, got imageInputCount=${String(imageRequest?.imageInputCount ?? 0)}`"
|
||||
detailsExpr: outbound.text
|
||||
```
|
||||
86
openclaw/qa/scenarios/media/native-image-generation.md
Normal file
86
openclaw/qa/scenarios/media/native-image-generation.md
Normal file
|
|
@ -0,0 +1,86 @@
|
|||
# Native image generation
|
||||
|
||||
```yaml qa-scenario
|
||||
id: native-image-generation
|
||||
title: Native image generation
|
||||
surface: image-generation
|
||||
coverage:
|
||||
primary:
|
||||
- media.image-generation
|
||||
secondary:
|
||||
- tools.native-image-generation
|
||||
objective: Verify image_generate appears when configured and returns a real saved media artifact.
|
||||
successCriteria:
|
||||
- image_generate appears in the effective tool inventory.
|
||||
- Agent triggers native image_generate.
|
||||
- Tool output returns a saved MEDIA path and the file exists.
|
||||
docsRefs:
|
||||
- docs/tools/image-generation.md
|
||||
- docs/providers/openai.md
|
||||
codeRefs:
|
||||
- src/agents/tools/image-generate-tool.ts
|
||||
- extensions/qa-lab/src/mock-openai-server.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify image_generate appears when configured and returns a real saved media artifact.
|
||||
config:
|
||||
prompt: "Image generation check: generate a QA lighthouse image and summarize it in one short sentence."
|
||||
promptSnippet: "Image generation check"
|
||||
generatedNeedle: "QA lighthouse"
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: enables image_generate and saves a real media artifact
|
||||
actions:
|
||||
- call: ensureImageGenerationConfigured
|
||||
args:
|
||||
- ref: env
|
||||
- call: createSession
|
||||
saveAs: sessionKey
|
||||
args:
|
||||
- ref: env
|
||||
- Image generation
|
||||
- call: readEffectiveTools
|
||||
saveAs: tools
|
||||
args:
|
||||
- ref: env
|
||||
- ref: sessionKey
|
||||
- assert:
|
||||
expr: "tools.has('image_generate')"
|
||||
message: image_generate not present after imageGenerationModel patch
|
||||
- call: reset
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:image-generate
|
||||
message:
|
||||
expr: config.prompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 45000)
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: outbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === 'qa-operator'"
|
||||
- expr: liveTurnTimeoutMs(env, 45000)
|
||||
- assert:
|
||||
expr: "!env.mock || ((await fetchJson(`${env.mock.baseUrl}/debug/requests`)).find((request) => String(request.allInputText ?? '').includes(config.promptSnippet))?.plannedToolName === 'image_generate')"
|
||||
message:
|
||||
expr: "`expected image_generate, got ${String((await fetchJson(`${env.mock.baseUrl}/debug/requests`)).find((request) => String(request.allInputText ?? '').includes(config.promptSnippet))?.plannedToolName ?? '')}`"
|
||||
- call: waitForCondition
|
||||
saveAs: generated
|
||||
args:
|
||||
- lambda:
|
||||
async: true
|
||||
expr: "!env.mock ? true : (await fetchJson(`${env.mock.baseUrl}/debug/image-generations`)).find((request) => request.model === 'gpt-image-1' && String(request.prompt ?? '').includes(config.generatedNeedle))"
|
||||
- 15000
|
||||
- 250
|
||||
- assert:
|
||||
expr: "!env.mock || Boolean(generated)"
|
||||
message:
|
||||
expr: "`image provider was never invoked`"
|
||||
detailsExpr: "env.mock ? `${outbound.text}\\nIMAGE_PROMPT:${generated.prompt ?? ''}` : outbound.text"
|
||||
```
|
||||
230
openclaw/qa/scenarios/memory/active-memory-preprompt-recall.md
Normal file
230
openclaw/qa/scenarios/memory/active-memory-preprompt-recall.md
Normal file
|
|
@ -0,0 +1,230 @@
|
|||
# Active Memory pre-reply recall
|
||||
|
||||
```yaml qa-scenario
|
||||
id: active-memory-preprompt-recall
|
||||
title: Active Memory pre-reply recall
|
||||
surface: memory
|
||||
coverage:
|
||||
primary:
|
||||
- memory.active-recall
|
||||
secondary:
|
||||
- memory.recall
|
||||
objective: Verify Active Memory surfaces a memory-only preference before the main reply, and that the same question stays unresolved when the plugin is off.
|
||||
plugins:
|
||||
- active-memory
|
||||
gatewayConfigPatch:
|
||||
plugins:
|
||||
entries:
|
||||
active-memory:
|
||||
enabled: true
|
||||
config:
|
||||
enabled: true
|
||||
agents:
|
||||
- qa
|
||||
allowedChatTypes:
|
||||
- direct
|
||||
logging: true
|
||||
persistTranscripts: true
|
||||
transcriptDir: qa-memory-e2e
|
||||
queryMode: recent
|
||||
maxSummaryChars: 220
|
||||
successCriteria:
|
||||
- With Active Memory off, the session shows no Active Memory plugin activity.
|
||||
- With Active Memory on, plugin-owned evidence shows the Active Memory sub-agent searched memory before the main reply.
|
||||
- Live lane proves the first user-visible reply uses the recalled preference.
|
||||
docsRefs:
|
||||
- docs/concepts/active-memory.md
|
||||
- docs/concepts/memory-search.md
|
||||
codeRefs:
|
||||
- extensions/active-memory/index.ts
|
||||
- extensions/qa-lab/src/suite.ts
|
||||
- extensions/qa-lab/src/mock-openai-server.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify Active Memory stays off when session-toggled off, runs memory search/get when enabled, and helps a live model answer with the recalled preference in the first visible reply.
|
||||
config:
|
||||
baselineConversationId: qa-active-memory-off
|
||||
activeConversationId: qa-active-memory-on
|
||||
memoryFact: "Stable QA movie night snack preference: lemon pepper wings with blue cheese."
|
||||
memoryQuery: "QA movie night snack lemon pepper wings blue cheese"
|
||||
expectedNeedle: lemon pepper wings
|
||||
prompt: "Silent snack recall check: what snack do I usually want for QA movie night? Reply in one short sentence."
|
||||
promptSnippet: "Silent snack recall check"
|
||||
transcriptDir: qa-memory-e2e
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: only active memory surfaces the hidden snack preference
|
||||
actions:
|
||||
- call: reset
|
||||
- call: fs.rm
|
||||
args:
|
||||
- expr: "path.join(env.gateway.workspaceDir, 'MEMORY.md')"
|
||||
- force: true
|
||||
- call: fs.rm
|
||||
args:
|
||||
- expr: "path.join(env.gateway.workspaceDir, 'memory', `${formatMemoryDreamingDay(Date.now())}.md`)"
|
||||
- force: true
|
||||
- call: fs.writeFile
|
||||
args:
|
||||
- expr: "path.join(env.gateway.workspaceDir, 'MEMORY.md')"
|
||||
- expr: "`${config.memoryFact}\\n`"
|
||||
- utf8
|
||||
- call: forceMemoryIndex
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
query:
|
||||
expr: config.memoryQuery
|
||||
expectedNeedle:
|
||||
expr: config.expectedNeedle
|
||||
- set: baselineSessionKey
|
||||
value:
|
||||
expr: "'agent:qa:qa-channel:direct:active-memory-off'"
|
||||
- set: activeSessionKey
|
||||
value:
|
||||
expr: "'agent:qa:qa-channel:direct:active-memory-on'"
|
||||
- set: transcriptRoot
|
||||
value:
|
||||
expr: "path.join(env.gateway.tempRoot, 'state', 'plugins', 'active-memory', 'transcripts', 'agents', 'qa', config.transcriptDir)"
|
||||
- set: toggleStorePath
|
||||
value:
|
||||
expr: "path.join(env.gateway.tempRoot, 'state', 'plugins', 'active-memory', 'session-toggles.json')"
|
||||
- call: fs.rm
|
||||
args:
|
||||
- ref: transcriptRoot
|
||||
- recursive: true
|
||||
force: true
|
||||
- call: fs.rm
|
||||
args:
|
||||
- ref: toggleStorePath
|
||||
- force: true
|
||||
- call: fs.mkdir
|
||||
args:
|
||||
- expr: "path.dirname(toggleStorePath)"
|
||||
- recursive: true
|
||||
- call: fs.writeFile
|
||||
args:
|
||||
- ref: toggleStorePath
|
||||
- expr: "`${JSON.stringify({ sessions: { [baselineSessionKey]: { disabled: true, updatedAt: Date.now() } } }, null, 2)}\\n`"
|
||||
- utf8
|
||||
- set: requestCountBeforeBaseline
|
||||
value:
|
||||
expr: "env.mock ? (await fetchJson(`${env.mock.baseUrl}/debug/requests`)).length : 0"
|
||||
- set: baselineStartIndex
|
||||
value:
|
||||
expr: "state.getSnapshot().messages.length"
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey:
|
||||
ref: baselineSessionKey
|
||||
message:
|
||||
expr: config.prompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 45000)
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: baselineOutbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === 'qa-operator'"
|
||||
- expr: liveTurnTimeoutMs(env, 30000)
|
||||
- sinceIndex:
|
||||
ref: baselineStartIndex
|
||||
- set: baselineLower
|
||||
value:
|
||||
expr: "normalizeLowercaseStringOrEmpty(baselineOutbound.text)"
|
||||
- if:
|
||||
expr: "Boolean(env.mock)"
|
||||
then:
|
||||
- set: baselineMockRequests
|
||||
value:
|
||||
expr: "(await fetchJson(`${env.mock.baseUrl}/debug/requests`)).slice(requestCountBeforeBaseline)"
|
||||
- set: baselineSessionStore
|
||||
value:
|
||||
expr: "await readRawQaSessionStore(env)"
|
||||
- assert:
|
||||
expr: "!Array.isArray(baselineSessionStore[baselineSessionKey]?.pluginDebugEntries) || !baselineSessionStore[baselineSessionKey].pluginDebugEntries.some((pluginEntry) => pluginEntry?.pluginId === 'active-memory')"
|
||||
message: baseline session unexpectedly recorded active-memory plugin activity
|
||||
- set: requestCountBeforeActive
|
||||
value:
|
||||
expr: "env.mock ? (await fetchJson(`${env.mock.baseUrl}/debug/requests`)).length : 0"
|
||||
- call: fs.writeFile
|
||||
args:
|
||||
- ref: toggleStorePath
|
||||
- expr: "'{}\\n'"
|
||||
- utf8
|
||||
- set: activeStartIndex
|
||||
value:
|
||||
expr: "state.getSnapshot().messages.length"
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey:
|
||||
ref: activeSessionKey
|
||||
message:
|
||||
expr: config.prompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 45000)
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: activeOutbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === 'qa-operator'"
|
||||
- expr: liveTurnTimeoutMs(env, 30000)
|
||||
- sinceIndex:
|
||||
ref: activeStartIndex
|
||||
- set: activeLower
|
||||
value:
|
||||
expr: "normalizeLowercaseStringOrEmpty(activeOutbound.text)"
|
||||
- if:
|
||||
expr: "!env.mock"
|
||||
then:
|
||||
- assert:
|
||||
expr: "activeLower.includes(normalizeLowercaseStringOrEmpty(config.expectedNeedle))"
|
||||
message:
|
||||
expr: "`active memory reply missed the hidden preference: ${activeOutbound.text}`"
|
||||
- call: waitForCondition
|
||||
saveAs: transcriptPath
|
||||
args:
|
||||
- lambda:
|
||||
async: true
|
||||
expr: "await (async () => { const entries = (await fs.readdir(transcriptRoot).catch(() => [])).filter((entry) => entry.endsWith('.jsonl')).toSorted(); return entries.length > 0 ? path.join(transcriptRoot, entries.at(-1)) : undefined; })()"
|
||||
- 10000
|
||||
- call: fs.readFile
|
||||
saveAs: transcriptText
|
||||
args:
|
||||
- ref: transcriptPath
|
||||
- utf8
|
||||
- assert:
|
||||
expr: "transcriptText.includes('memory_search')"
|
||||
message: active memory transcript missing memory_search
|
||||
- assert:
|
||||
expr: "transcriptText.includes('memory_get')"
|
||||
message: active memory transcript missing memory_get
|
||||
- call: waitForCondition
|
||||
saveAs: activeSessionEntry
|
||||
args:
|
||||
- lambda:
|
||||
async: true
|
||||
expr: "await (async () => { const store = await readRawQaSessionStore(env); const entry = store[activeSessionKey]; if (!entry || !Array.isArray(entry.pluginDebugEntries)) return undefined; return entry.pluginDebugEntries.some((pluginEntry) => pluginEntry?.pluginId === 'active-memory' && Array.isArray(pluginEntry.lines) && pluginEntry.lines.some((line) => line.includes('Active Memory: status=ok'))) ? entry : undefined; })()"
|
||||
- 10000
|
||||
- if:
|
||||
expr: "Boolean(env.mock)"
|
||||
then:
|
||||
- set: mockRequests
|
||||
value:
|
||||
expr: "(await fetchJson(`${env.mock.baseUrl}/debug/requests`)).slice(requestCountBeforeActive)"
|
||||
- assert:
|
||||
expr: "mockRequests.some((request) => request.allInputText.includes('You are a memory search agent.') && request.plannedToolName === 'memory_search')"
|
||||
message: expected mock Active Memory search request
|
||||
- assert:
|
||||
expr: "mockRequests.some((request) => request.allInputText.includes('You are a memory search agent.') && request.plannedToolName === 'memory_get')"
|
||||
message: expected mock Active Memory memory_get request
|
||||
detailsExpr: "`${activeOutbound.text}\\n\\ntranscript=${transcriptPath}`"
|
||||
```
|
||||
285
openclaw/qa/scenarios/memory/memory-dreaming-sweep.md
Normal file
285
openclaw/qa/scenarios/memory/memory-dreaming-sweep.md
Normal file
|
|
@ -0,0 +1,285 @@
|
|||
# Memory dreaming sweep
|
||||
|
||||
```yaml qa-scenario
|
||||
id: memory-dreaming-sweep
|
||||
title: Memory dreaming sweep
|
||||
surface: memory
|
||||
coverage:
|
||||
primary:
|
||||
- memory.dreaming
|
||||
objective: Verify enabling dreaming creates the managed sweep, stages light and REM artifacts, and consolidates repeated recall signals into durable memory.
|
||||
successCriteria:
|
||||
- Dreaming can be enabled and doctor.memory.status reports the managed sweep cron.
|
||||
- Repeated recall signals give the dreaming sweep real material to process.
|
||||
- A dreaming sweep writes Light Sleep and REM Sleep blocks, then promotes the canary into MEMORY.md.
|
||||
docsRefs:
|
||||
- docs/concepts/dreaming.md
|
||||
- docs/reference/memory-config.md
|
||||
- docs/web/control-ui.md
|
||||
codeRefs:
|
||||
- extensions/memory-core/src/dreaming.ts
|
||||
- extensions/memory-core/src/dreaming-phases.ts
|
||||
- src/gateway/server-methods/doctor.ts
|
||||
- extensions/qa-lab/src/suite.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify enabling dreaming creates the managed sweep, stages light and REM artifacts, and consolidates repeated recall signals into durable memory.
|
||||
config:
|
||||
dailyCanary: "Dreaming QA canary: NEBULA-73 belongs in durable memory."
|
||||
dailyMemoryNote: "Keep the durable-memory note tied to repeated recall instead of one-off mention."
|
||||
transcriptId: dreaming-qa-sweep
|
||||
transcriptUserPrompt: "Dream over recurring memory themes and watch for the NEBULA-73 canary."
|
||||
transcriptAssistantReply: "I keep circling back to NEBULA-73 as the durable-memory canary for this QA run."
|
||||
searchQueries:
|
||||
- "dreaming qa canary nebula-73"
|
||||
- "durable memory canary nebula 73"
|
||||
- "which canary belongs to the dreaming qa check"
|
||||
expectedNeedle: "NEBULA-73"
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: enables dreaming and registers the managed sweep cron
|
||||
actions:
|
||||
- call: readConfigSnapshot
|
||||
saveAs: original
|
||||
args:
|
||||
- ref: env
|
||||
- set: pluginEntries
|
||||
value:
|
||||
expr: "original.config.plugins && typeof original.config.plugins === 'object' ? original.config.plugins.entries : undefined"
|
||||
- set: memoryCoreEntry
|
||||
value:
|
||||
expr: "pluginEntries && typeof pluginEntries['memory-core'] === 'object' ? pluginEntries['memory-core'] : undefined"
|
||||
- set: memoryCoreConfig
|
||||
value:
|
||||
expr: "memoryCoreEntry && typeof memoryCoreEntry.config === 'object' ? memoryCoreEntry.config : undefined"
|
||||
- set: originalDreaming
|
||||
value:
|
||||
expr: "memoryCoreConfig?.dreaming"
|
||||
- call: patchConfig
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
patch:
|
||||
plugins:
|
||||
entries:
|
||||
memory-core:
|
||||
config:
|
||||
dreaming:
|
||||
enabled: true
|
||||
phases:
|
||||
deep:
|
||||
minScore: 0
|
||||
minRecallCount: 3
|
||||
minUniqueQueries: 3
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- call: waitForQaChannelReady
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- try:
|
||||
actions:
|
||||
- call: waitForCondition
|
||||
saveAs: status
|
||||
args:
|
||||
- lambda:
|
||||
async: true
|
||||
expr: "(() => readDoctorMemoryStatus(env).then((payload) => payload.dreaming?.phases?.deep?.managedCronPresent === true ? payload : undefined))()"
|
||||
- 30000
|
||||
- 500
|
||||
- call: listCronJobs
|
||||
saveAs: jobs
|
||||
args:
|
||||
- ref: env
|
||||
- set: managed
|
||||
value:
|
||||
expr: "jobs.find((job) => job.name === 'Memory Dreaming Promotion' && job.payload?.kind === 'systemEvent' && job.payload.text === '__openclaw_memory_core_short_term_promotion_dream__')"
|
||||
- assert:
|
||||
expr: "Boolean(managed?.id)"
|
||||
message: managed dreaming cron job missing after enablement
|
||||
- set: dreamingOriginal
|
||||
value:
|
||||
expr: "structuredClone(originalDreaming)"
|
||||
- set: dreamingCronId
|
||||
value:
|
||||
expr: "managed.id"
|
||||
catchAs: enableError
|
||||
catch:
|
||||
- call: patchConfig
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
patch:
|
||||
plugins:
|
||||
entries:
|
||||
memory-core:
|
||||
config:
|
||||
dreaming:
|
||||
expr: "originalDreaming === undefined ? null : structuredClone(originalDreaming)"
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- call: waitForQaChannelReady
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- throw:
|
||||
expr: enableError
|
||||
detailsExpr: "JSON.stringify({ enabled: status.dreaming?.enabled ?? false, managedCronPresent: status.dreaming?.phases?.deep?.managedCronPresent ?? false, nextRunAtMs: status.dreaming?.phases?.deep?.nextRunAtMs ?? null })"
|
||||
|
||||
- name: runs the sweep after repeated recall signals and writes promotion artifacts
|
||||
actions:
|
||||
- assert:
|
||||
expr: "Boolean(dreamingCronId)"
|
||||
message: missing managed dreaming cron id
|
||||
- set: cronId
|
||||
value:
|
||||
ref: dreamingCronId
|
||||
- set: dreamingDay
|
||||
value:
|
||||
expr: "formatMemoryDreamingDay(Date.now())"
|
||||
- set: dailyPath
|
||||
value:
|
||||
expr: "path.join(env.gateway.workspaceDir, 'memory', `${dreamingDay}.md`)"
|
||||
- set: lightReportPath
|
||||
value:
|
||||
expr: "path.join(env.gateway.workspaceDir, 'memory', 'dreaming', 'light', `${dreamingDay}.md`)"
|
||||
- set: remReportPath
|
||||
value:
|
||||
expr: "path.join(env.gateway.workspaceDir, 'memory', 'dreaming', 'rem', `${dreamingDay}.md`)"
|
||||
- set: memoryPath
|
||||
value:
|
||||
expr: "path.join(env.gateway.workspaceDir, 'MEMORY.md')"
|
||||
- set: homeDir
|
||||
value:
|
||||
expr: "env.gateway.runtimeEnv.HOME ?? env.gateway.runtimeEnv.OPENCLAW_HOME ?? env.gateway.tempRoot"
|
||||
- set: sessionsDir
|
||||
value:
|
||||
expr: "resolveSessionTranscriptsDirForAgent('qa', env.gateway.runtimeEnv, () => homeDir)"
|
||||
- set: transcriptPath
|
||||
value:
|
||||
expr: "path.join(sessionsDir, `${config.transcriptId}.jsonl`)"
|
||||
- try:
|
||||
actions:
|
||||
- call: fs.mkdir
|
||||
args:
|
||||
- expr: "path.dirname(dailyPath)"
|
||||
- recursive: true
|
||||
- call: fs.mkdir
|
||||
args:
|
||||
- ref: sessionsDir
|
||||
- recursive: true
|
||||
- call: fs.writeFile
|
||||
args:
|
||||
- ref: dailyPath
|
||||
- expr: "[`# ${dreamingDay}`, '', `- ${config.dailyCanary}`, `- ${config.dailyMemoryNote}`].join('\\n') + '\\n'"
|
||||
- utf8
|
||||
- set: now
|
||||
value:
|
||||
expr: "Date.now()"
|
||||
- call: fs.writeFile
|
||||
args:
|
||||
- ref: transcriptPath
|
||||
- expr: "[JSON.stringify({ type: 'session', id: config.transcriptId, timestamp: new Date(now - 120000).toISOString() }), JSON.stringify({ type: 'message', message: { role: 'user', timestamp: new Date(now - 90000).toISOString(), content: [{ type: 'text', text: config.transcriptUserPrompt }] } }), JSON.stringify({ type: 'message', message: { role: 'assistant', timestamp: new Date(now - 60000).toISOString(), content: [{ type: 'text', text: config.transcriptAssistantReply }] } })].join('\\n') + '\\n'"
|
||||
- utf8
|
||||
- call: fs.rm
|
||||
args:
|
||||
- ref: memoryPath
|
||||
- force: true
|
||||
- call: forceMemoryIndex
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
query:
|
||||
expr: "config.searchQueries[0]"
|
||||
expectedNeedle:
|
||||
expr: config.expectedNeedle
|
||||
- call: sleep
|
||||
args:
|
||||
- 1000
|
||||
- forEach:
|
||||
items:
|
||||
expr: config.searchQueries
|
||||
item: query
|
||||
actions:
|
||||
- call: runQaCli
|
||||
saveAs: payload
|
||||
args:
|
||||
- ref: env
|
||||
- - memory
|
||||
- search
|
||||
- --agent
|
||||
- qa
|
||||
- --json
|
||||
- --query
|
||||
- ref: query
|
||||
- timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 60000)
|
||||
json: true
|
||||
- assert:
|
||||
expr: "JSON.stringify(payload.results ?? []).includes(config.expectedNeedle)"
|
||||
message:
|
||||
expr: "`memory search missed dreaming canary for query: ${query}`"
|
||||
- set: cronRunStartedAt
|
||||
value:
|
||||
expr: "Date.now()"
|
||||
- call: env.gateway.call
|
||||
saveAs: cronRun
|
||||
args:
|
||||
- cron.run
|
||||
- id:
|
||||
ref: cronId
|
||||
mode: force
|
||||
- timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 30000)
|
||||
- assert:
|
||||
expr: "cronRun.enqueued === true && Boolean(cronRun.runId)"
|
||||
message:
|
||||
expr: "`dreaming cron did not enqueue a background run: ${JSON.stringify(cronRun)}`"
|
||||
- call: waitForCronRunCompletion
|
||||
saveAs: finishedRun
|
||||
args:
|
||||
- callGateway:
|
||||
expr: "(method, rpcParams, opts) => env.gateway.call(method, rpcParams, opts)"
|
||||
jobId:
|
||||
ref: cronId
|
||||
afterTs:
|
||||
ref: cronRunStartedAt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 90000)
|
||||
- assert:
|
||||
expr: "finishedRun.status === 'ok'"
|
||||
message:
|
||||
expr: "`dreaming cron finished with ${finishedRun.status ?? 'unknown'}: ${JSON.stringify(finishedRun)}`"
|
||||
- call: waitForCondition
|
||||
saveAs: promoted
|
||||
args:
|
||||
- lambda:
|
||||
async: true
|
||||
expr: "(async () => { const status = await readDoctorMemoryStatus(env); const lightReport = await fs.readFile(lightReportPath, 'utf8').catch(() => ''); const remReport = await fs.readFile(remReportPath, 'utf8').catch(() => ''); const promotedMemory = await fs.readFile(memoryPath, 'utf8').catch(() => ''); if (!lightReport.includes('# Light Sleep')) return undefined; if (!remReport.includes('# REM Sleep')) return undefined; if (!promotedMemory.includes(config.expectedNeedle)) return undefined; if (status.dreaming?.phases?.deep?.managedCronPresent !== true) return undefined; if ((status.dreaming?.promotedTotal ?? 0) < 1) return undefined; return { status, lightReport, remReport, promotedMemory }; })()"
|
||||
- expr: liveTurnTimeoutMs(env, 90000)
|
||||
- 1000
|
||||
finally:
|
||||
- call: patchConfig
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
patch:
|
||||
plugins:
|
||||
entries:
|
||||
memory-core:
|
||||
config:
|
||||
dreaming:
|
||||
expr: "dreamingOriginal === undefined ? null : structuredClone(dreamingOriginal)"
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- call: waitForQaChannelReady
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
detailsExpr: "JSON.stringify({ promotedTotal: promoted.status.dreaming?.promotedTotal ?? 0, shortTermCount: promoted.status.dreaming?.shortTermCount ?? 0, phaseSignalCount: promoted.status.dreaming?.phaseSignalCount ?? 0, lightSleep: promoted.lightReport.includes('# Light Sleep'), remSleep: promoted.remReport.includes('# REM Sleep') })"
|
||||
```
|
||||
164
openclaw/qa/scenarios/memory/memory-failure-fallback.md
Normal file
164
openclaw/qa/scenarios/memory/memory-failure-fallback.md
Normal file
|
|
@ -0,0 +1,164 @@
|
|||
# Memory failure fallback
|
||||
|
||||
```yaml qa-scenario
|
||||
id: memory-failure-fallback
|
||||
title: Memory failure fallback
|
||||
surface: memory
|
||||
coverage:
|
||||
primary:
|
||||
- memory.failure-handling
|
||||
secondary:
|
||||
- runtime.fallbacks
|
||||
objective: Verify the agent degrades gracefully when memory tools are unavailable and the answer exists only in memory-backed notes.
|
||||
successCriteria:
|
||||
- Memory tools are absent from the effective tool inventory.
|
||||
- Agent does not hallucinate the hidden fact.
|
||||
- Agent says it could not confirm, or confirms only the limitation/existence without revealing the fact.
|
||||
docsRefs:
|
||||
- docs/concepts/memory.md
|
||||
- docs/tools/index.md
|
||||
codeRefs:
|
||||
- extensions/memory-core/src/tools.ts
|
||||
- extensions/qa-lab/src/suite.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify the agent degrades gracefully when memory tools are unavailable and the answer exists only in memory-backed notes.
|
||||
config:
|
||||
memoryFact: "Do not reveal directly: fallback fact is ORBIT-9."
|
||||
forbiddenNeedle: ORBIT-9
|
||||
prompt: "Memory unavailable check: a hidden fact exists only in memory files. If you cannot confirm it, say so clearly and do not guess."
|
||||
gracefulFallbackAny:
|
||||
- could not confirm
|
||||
- can't confirm
|
||||
- can’t confirm
|
||||
- cannot confirm
|
||||
- i can confirm there is a hidden fact
|
||||
- will not guess
|
||||
- won't guess
|
||||
- won’t guess
|
||||
- should not guess
|
||||
- cannot see
|
||||
- can't see
|
||||
- can’t see
|
||||
- should not reveal
|
||||
- won't reveal
|
||||
- won’t reveal
|
||||
- will not reveal
|
||||
- "confirmed: the hidden fact is present"
|
||||
- hidden fact is present
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: falls back cleanly when group:memory tools are denied
|
||||
actions:
|
||||
- call: readConfigSnapshot
|
||||
saveAs: original
|
||||
args:
|
||||
- ref: env
|
||||
- set: originalTools
|
||||
value:
|
||||
expr: "original.config.tools && typeof original.config.tools === 'object' ? original.config.tools : null"
|
||||
- set: originalToolsDeny
|
||||
value:
|
||||
expr: "originalTools ? (Object.prototype.hasOwnProperty.call(originalTools, 'deny') ? structuredClone(originalTools.deny) : undefined) : undefined"
|
||||
- call: fs.writeFile
|
||||
args:
|
||||
- expr: "path.join(env.gateway.workspaceDir, 'MEMORY.md')"
|
||||
- expr: "`${config.memoryFact}\\n`"
|
||||
- utf8
|
||||
- set: deniedTools
|
||||
value:
|
||||
expr: "Array.isArray(originalToolsDeny) ? originalToolsDeny.map((entry) => String(entry)) : []"
|
||||
- set: nextDeniedTools
|
||||
value:
|
||||
expr: "deniedTools.concat(['group:memory', 'read']).filter((value, index, array) => array.indexOf(value) === index)"
|
||||
- call: patchConfig
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
patch:
|
||||
tools:
|
||||
deny:
|
||||
ref: nextDeniedTools
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- call: waitForQaChannelReady
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- try:
|
||||
actions:
|
||||
- call: createSession
|
||||
saveAs: sessionKey
|
||||
args:
|
||||
- ref: env
|
||||
- Memory fallback
|
||||
- call: readEffectiveTools
|
||||
saveAs: tools
|
||||
args:
|
||||
- ref: env
|
||||
- ref: sessionKey
|
||||
- assert:
|
||||
expr: "!tools.has('memory_search') && !tools.has('memory_get') && !tools.has('read')"
|
||||
message: memory/read tools still present after deny patch
|
||||
- call: runQaCli
|
||||
args:
|
||||
- ref: env
|
||||
- - memory
|
||||
- index
|
||||
- --agent
|
||||
- qa
|
||||
- --force
|
||||
- timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 60000)
|
||||
- call: reset
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:memory-failure
|
||||
message:
|
||||
expr: config.prompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 30000)
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: outbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === 'qa-operator'"
|
||||
- expr: liveTurnTimeoutMs(env, 30000)
|
||||
- set: lower
|
||||
value:
|
||||
expr: "normalizeLowercaseStringOrEmpty(outbound.text)"
|
||||
- assert:
|
||||
expr: "!outbound.text.includes(config.forbiddenNeedle)"
|
||||
message:
|
||||
expr: "`hallucinated hidden fact: ${outbound.text}`"
|
||||
- set: gracefulFallback
|
||||
value:
|
||||
expr: "config.gracefulFallbackAny.some((needle) => lower.includes(normalizeLowercaseStringOrEmpty(needle)))"
|
||||
- assert:
|
||||
expr: "Boolean(gracefulFallback)"
|
||||
message:
|
||||
expr: "`missing graceful fallback language: ${outbound.text}`"
|
||||
finally:
|
||||
- call: patchConfig
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
patch:
|
||||
tools:
|
||||
deny:
|
||||
expr: "originalToolsDeny === undefined ? null : originalToolsDeny"
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- call: waitForQaChannelReady
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
detailsExpr: outbound.text
|
||||
```
|
||||
117
openclaw/qa/scenarios/memory/memory-recall.md
Normal file
117
openclaw/qa/scenarios/memory/memory-recall.md
Normal file
|
|
@ -0,0 +1,117 @@
|
|||
# Memory recall after context switch
|
||||
|
||||
<!--
|
||||
This scenario deliberately stays prose-only and does NOT gate on a
|
||||
`/debug/requests` tool-call assertion, even though it is one of the
|
||||
scenarios in the parity pack. The adversarial review in the umbrella
|
||||
#64227 thread called this out as a coverage gap, but the underlying
|
||||
behavior the scenario tests is legitimately prose-shaped: the agent is
|
||||
supposed to pull a prior-turn fact ("ALPHA-7") back across an
|
||||
intervening context switch and reply with the code. In a real
|
||||
conversation, the model can do this EITHER by calling a memory-search
|
||||
tool (which the qa-lab mock server doesn't currently expose) OR by
|
||||
reading the fact directly from prior-turn context in its own
|
||||
conversation window. Both strategies are valid parity behavior.
|
||||
|
||||
Forcing a `plannedToolName` assertion here would either require
|
||||
extending the mock with a synthetic `memory_search` tool lane (PR O
|
||||
scope, not PR J) or fabricating a tool-call requirement the real
|
||||
providers never implement. Either path would make this scenario test
|
||||
the harness, not the models. So we keep it prose-only, covered by the
|
||||
`recallExpectedAny` / `rememberAckAny` assertions above, and flag the
|
||||
exception explicitly rather than silently.
|
||||
|
||||
Criterion 2 of the parity completion gate (no fake progress or fake
|
||||
tool completion) is enforced for this scenario through the parity
|
||||
report's failure-tone fake-success detector: a scenario marked `pass`
|
||||
whose details text matches patterns like "timed out", "failed to",
|
||||
"could not" gets flagged via `SUSPICIOUS_PASS_FAILURE_TONE_PATTERNS`
|
||||
in `extensions/qa-lab/src/agentic-parity-report.ts`. Positive-tone
|
||||
detection was removed because it false-positives on legitimate passes
|
||||
where the details field is the model's outbound prose.
|
||||
-->
|
||||
|
||||
```yaml qa-scenario
|
||||
id: memory-recall
|
||||
title: Memory recall after context switch
|
||||
surface: memory
|
||||
coverage:
|
||||
primary:
|
||||
- memory.recall
|
||||
objective: Verify the agent can store a fact, switch topics, then recall the fact accurately later.
|
||||
successCriteria:
|
||||
- Agent acknowledges the seeded fact.
|
||||
- Agent later recalls the same fact correctly.
|
||||
- Recall stays scoped to the active QA conversation.
|
||||
docsRefs:
|
||||
- docs/help/testing.md
|
||||
codeRefs:
|
||||
- extensions/qa-lab/src/scenario.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify the agent can store a fact, switch topics, then recall the fact accurately later.
|
||||
config:
|
||||
resetDurableMemory: true
|
||||
rememberPrompt: "Please remember this fact for later: the QA canary code is ALPHA-7. Use your normal memory mechanism, avoid manual repo cleanup, and reply exactly `Remembered ALPHA-7.` once stored."
|
||||
rememberAckAny:
|
||||
- remembered alpha-7
|
||||
recallPrompt: "What was the QA canary code I asked you to remember earlier? Reply with the code only, plus at most one short sentence."
|
||||
recallExpectedAny:
|
||||
- alpha-7
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: stores the canary fact
|
||||
actions:
|
||||
- assert:
|
||||
expr: "!config.resetDurableMemory || true"
|
||||
- call: fs.rm
|
||||
args:
|
||||
- expr: "path.join(env.gateway.workspaceDir, 'MEMORY.md')"
|
||||
- force: true
|
||||
- call: fs.rm
|
||||
args:
|
||||
- expr: "path.join(env.gateway.workspaceDir, 'memory', `${formatMemoryDreamingDay(Date.now())}.md`)"
|
||||
- force: true
|
||||
- call: reset
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:memory
|
||||
message:
|
||||
expr: config.rememberPrompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 60000)
|
||||
- set: rememberAckAny
|
||||
value:
|
||||
expr: config.rememberAckAny.map(normalizeLowercaseStringOrEmpty)
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: outbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === 'qa-operator' && rememberAckAny.some((needle) => normalizeLowercaseStringOrEmpty(candidate.text).includes(needle))"
|
||||
detailsExpr: outbound.text
|
||||
- name: recalls the same fact later
|
||||
actions:
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:memory
|
||||
message:
|
||||
expr: config.recallPrompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 60000)
|
||||
- set: recallExpectedAny
|
||||
value:
|
||||
expr: config.recallExpectedAny.map(normalizeLowercaseStringOrEmpty)
|
||||
- call: waitForCondition
|
||||
saveAs: outbound
|
||||
args:
|
||||
- lambda:
|
||||
expr: "state.getSnapshot().messages.filter((candidate) => candidate.direction === 'outbound' && candidate.conversation.id === 'qa-operator' && recallExpectedAny.some((needle) => normalizeLowercaseStringOrEmpty(candidate.text).includes(needle))).at(-1)"
|
||||
- 20000
|
||||
detailsExpr: outbound.text
|
||||
```
|
||||
89
openclaw/qa/scenarios/memory/memory-tools-channel-context.md
Normal file
89
openclaw/qa/scenarios/memory/memory-tools-channel-context.md
Normal file
|
|
@ -0,0 +1,89 @@
|
|||
# Memory tools in channel context
|
||||
|
||||
```yaml qa-scenario
|
||||
id: memory-tools-channel-context
|
||||
title: Memory tools in channel context
|
||||
surface: memory
|
||||
coverage:
|
||||
primary:
|
||||
- memory.tools
|
||||
secondary:
|
||||
- channels.group-messages
|
||||
objective: Verify the agent uses memory_search and memory_get in a shared channel when the answer lives only in memory files, not the live transcript.
|
||||
successCriteria:
|
||||
- Agent uses memory_search before answering.
|
||||
- Agent narrows with memory_get before answering.
|
||||
- Final reply returns the memory-only fact correctly in-channel.
|
||||
docsRefs:
|
||||
- docs/concepts/memory.md
|
||||
- docs/concepts/memory-search.md
|
||||
codeRefs:
|
||||
- extensions/memory-core/src/tools.ts
|
||||
- extensions/qa-lab/src/suite.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify the agent uses memory_search and memory_get in a shared channel when the answer lives only in memory files, not the live transcript.
|
||||
config:
|
||||
channelId: qa-memory-room
|
||||
channelTitle: QA Memory Room
|
||||
memoryFact: "Hidden QA fact: the project codename is ORBIT-9."
|
||||
memoryQuery: "project codename ORBIT-9"
|
||||
expectedNeedle: ORBIT-9
|
||||
prompt: "@openclaw Memory tools check: what is the hidden project codename stored only in memory? Use memory tools first."
|
||||
promptSnippet: "Memory tools check"
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: uses memory_search plus memory_get before answering in-channel
|
||||
actions:
|
||||
- call: reset
|
||||
- call: fs.writeFile
|
||||
args:
|
||||
- expr: "path.join(env.gateway.workspaceDir, 'MEMORY.md')"
|
||||
- expr: "`${config.memoryFact}\\n`"
|
||||
- utf8
|
||||
- call: forceMemoryIndex
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
query:
|
||||
expr: config.memoryQuery
|
||||
expectedNeedle:
|
||||
expr: config.expectedNeedle
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- call: waitForQaChannelReady
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- call: state.addInboundMessage
|
||||
args:
|
||||
- conversation:
|
||||
id:
|
||||
expr: config.channelId
|
||||
kind: channel
|
||||
title:
|
||||
expr: config.channelTitle
|
||||
senderId: alice
|
||||
senderName: Alice
|
||||
text:
|
||||
expr: config.prompt
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: outbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === config.channelId && candidate.text.includes(config.expectedNeedle)"
|
||||
- expr: liveTurnTimeoutMs(env, 30000)
|
||||
- assert:
|
||||
expr: "!env.mock || (await fetchJson(`${env.mock.baseUrl}/debug/requests`)).filter((request) => String(request.allInputText ?? '').includes(config.promptSnippet)).some((request) => request.plannedToolName === 'memory_search')"
|
||||
message: expected memory_search in mock request plan
|
||||
- assert:
|
||||
expr: "!env.mock || (await fetchJson(`${env.mock.baseUrl}/debug/requests`)).some((request) => request.plannedToolName === 'memory_get')"
|
||||
message: expected memory_get in mock request plan
|
||||
detailsExpr: outbound.text
|
||||
```
|
||||
190
openclaw/qa/scenarios/memory/session-memory-ranking.md
Normal file
190
openclaw/qa/scenarios/memory/session-memory-ranking.md
Normal file
|
|
@ -0,0 +1,190 @@
|
|||
# Session memory ranking
|
||||
|
||||
```yaml qa-scenario
|
||||
id: session-memory-ranking
|
||||
title: Session memory ranking
|
||||
surface: memory
|
||||
coverage:
|
||||
primary:
|
||||
- memory.ranking
|
||||
secondary:
|
||||
- memory.recall
|
||||
objective: Verify session-transcript memory can outrank stale durable notes and drive the final answer toward the newer fact.
|
||||
successCriteria:
|
||||
- Session memory indexing is enabled for the scenario.
|
||||
- Search ranks the newer transcript-backed fact ahead of the stale durable note.
|
||||
- The agent uses memory tools and answers with the current fact, not the stale one.
|
||||
docsRefs:
|
||||
- docs/concepts/memory-search.md
|
||||
- docs/reference/memory-config.md
|
||||
codeRefs:
|
||||
- extensions/memory-core/src/tools.ts
|
||||
- extensions/memory-core/src/memory/manager.ts
|
||||
- extensions/qa-lab/src/suite.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify session-transcript memory can outrank stale durable notes and drive the final answer toward the newer fact.
|
||||
config:
|
||||
staleFact: ORBIT-9
|
||||
currentFact: ORBIT-10
|
||||
transcriptId: qa-session-memory-ranking
|
||||
transcriptQuestion: "What is the current Project Nebula codename?"
|
||||
transcriptAnswer: "The current Project Nebula codename is ORBIT-10."
|
||||
prompt: "Session memory ranking check: what is the current Project Nebula codename? Use memory tools first. If durable notes conflict with newer indexed session transcripts, prefer the newer current fact."
|
||||
promptSnippet: "Session memory ranking check"
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: prefers the newer transcript-backed fact over the stale durable note
|
||||
actions:
|
||||
- set: staleFact
|
||||
value:
|
||||
expr: config.staleFact
|
||||
- set: currentFact
|
||||
value:
|
||||
expr: config.currentFact
|
||||
- call: readConfigSnapshot
|
||||
saveAs: original
|
||||
args:
|
||||
- ref: env
|
||||
- set: originalMemorySearch
|
||||
value:
|
||||
expr: "original.config.agents && typeof original.config.agents === 'object' && typeof original.config.agents.defaults === 'object' ? original.config.agents.defaults.memorySearch : undefined"
|
||||
- call: patchConfig
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
patch:
|
||||
agents:
|
||||
defaults:
|
||||
memorySearch:
|
||||
sources:
|
||||
- memory
|
||||
- sessions
|
||||
experimental:
|
||||
sessionMemory: true
|
||||
query:
|
||||
minScore: 0
|
||||
hybrid:
|
||||
enabled: true
|
||||
temporalDecay:
|
||||
enabled: true
|
||||
halfLifeDays: 1
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- call: waitForQaChannelReady
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- try:
|
||||
actions:
|
||||
- set: memoryDir
|
||||
value:
|
||||
expr: "path.join(env.gateway.workspaceDir, 'memory')"
|
||||
- call: fs.mkdir
|
||||
args:
|
||||
- ref: memoryDir
|
||||
- recursive: true
|
||||
- set: staleMemoryPath
|
||||
value:
|
||||
expr: "path.join(memoryDir, '2020-01-01.md')"
|
||||
- call: fs.writeFile
|
||||
args:
|
||||
- ref: staleMemoryPath
|
||||
- expr: "`${'Project Nebula stale codename: '}${staleFact}.\\n`"
|
||||
- utf8
|
||||
- set: staleAt
|
||||
value:
|
||||
expr: "new Date('2020-01-01T00:00:00.000Z')"
|
||||
- call: fs.utimes
|
||||
args:
|
||||
- ref: staleMemoryPath
|
||||
- ref: staleAt
|
||||
- ref: staleAt
|
||||
- set: transcriptsDir
|
||||
value:
|
||||
expr: "resolveSessionTranscriptsDirForAgent('qa', env.gateway.runtimeEnv, () => env.gateway.runtimeEnv.HOME ?? path.join(env.gateway.tempRoot, 'home'))"
|
||||
- call: fs.mkdir
|
||||
args:
|
||||
- ref: transcriptsDir
|
||||
- recursive: true
|
||||
- set: transcriptPath
|
||||
value:
|
||||
expr: "path.join(transcriptsDir, `${config.transcriptId}.jsonl`)"
|
||||
- set: now
|
||||
value:
|
||||
expr: "Date.now()"
|
||||
- call: fs.writeFile
|
||||
args:
|
||||
- ref: transcriptPath
|
||||
- expr: "[JSON.stringify({ type: 'session', id: config.transcriptId, timestamp: new Date(now - 120000).toISOString() }), JSON.stringify({ type: 'message', message: { role: 'user', timestamp: new Date(now - 90000).toISOString(), content: [{ type: 'text', text: config.transcriptQuestion }] } }), JSON.stringify({ type: 'message', message: { role: 'assistant', timestamp: new Date(now - 60000).toISOString(), content: [{ type: 'text', text: config.transcriptAnswer }] } })].join('\\n') + '\\n'"
|
||||
- utf8
|
||||
- call: forceMemoryIndex
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
query:
|
||||
expr: "`current Project Nebula codename ${currentFact}`"
|
||||
expectedNeedle:
|
||||
ref: currentFact
|
||||
- call: reset
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:session-memory-ranking
|
||||
message:
|
||||
expr: config.prompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 45000)
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: outbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === 'qa-operator' && candidate.text.includes(currentFact)"
|
||||
- expr: liveTurnTimeoutMs(env, 45000)
|
||||
- set: lower
|
||||
value:
|
||||
expr: "normalizeLowercaseStringOrEmpty(outbound.text)"
|
||||
- set: staleLeak
|
||||
value:
|
||||
expr: "outbound.text.includes(staleFact) && !lower.includes('stale') && !lower.includes('older') && !lower.includes('previous')"
|
||||
- assert:
|
||||
expr: "!staleLeak"
|
||||
message:
|
||||
expr: "`stale durable fact leaked through: ${outbound.text}`"
|
||||
- if:
|
||||
expr: "Boolean(env.mock)"
|
||||
then:
|
||||
- call: fetchJson
|
||||
saveAs: requests
|
||||
args:
|
||||
- expr: "`${env.mock.baseUrl}/debug/requests`"
|
||||
- set: relevant
|
||||
value:
|
||||
expr: "requests.filter((request) => String(request.allInputText ?? '').includes(config.promptSnippet))"
|
||||
- assert:
|
||||
expr: "relevant.some((request) => request.plannedToolName === 'memory_search')"
|
||||
message: expected memory_search in session memory ranking flow
|
||||
finally:
|
||||
- call: patchConfig
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
patch:
|
||||
agents:
|
||||
defaults:
|
||||
memorySearch:
|
||||
expr: "originalMemorySearch === undefined ? null : structuredClone(originalMemorySearch)"
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- call: waitForQaChannelReady
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
detailsExpr: outbound.text
|
||||
```
|
||||
108
openclaw/qa/scenarios/memory/thread-memory-isolation.md
Normal file
108
openclaw/qa/scenarios/memory/thread-memory-isolation.md
Normal file
|
|
@ -0,0 +1,108 @@
|
|||
# Thread memory isolation
|
||||
|
||||
```yaml qa-scenario
|
||||
id: thread-memory-isolation
|
||||
title: Thread memory isolation
|
||||
surface: memory
|
||||
coverage:
|
||||
primary:
|
||||
- memory.thread-isolation
|
||||
secondary:
|
||||
- channels.threads
|
||||
objective: Verify a memory-backed answer requested inside a thread stays in-thread and does not leak into the root channel.
|
||||
successCriteria:
|
||||
- Agent uses memory tools inside the thread.
|
||||
- The hidden fact is answered correctly in the thread.
|
||||
- No root-channel outbound message leaks during the threaded memory reply.
|
||||
docsRefs:
|
||||
- docs/concepts/memory-search.md
|
||||
- docs/channels/qa-channel.md
|
||||
- docs/channels/group-messages.md
|
||||
codeRefs:
|
||||
- extensions/memory-core/src/tools.ts
|
||||
- extensions/qa-channel/src/protocol.ts
|
||||
- extensions/qa-lab/src/suite.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify a memory-backed answer requested inside a thread stays in-thread and does not leak into the root channel.
|
||||
config:
|
||||
memoryFact: "Thread-hidden codename: ORBIT-22."
|
||||
memoryQuery: "hidden thread codename ORBIT-22"
|
||||
expectedNeedle: "ORBIT-22"
|
||||
channelId: qa-room
|
||||
channelTitle: QA Room
|
||||
threadTitle: "Thread memory QA"
|
||||
prompt: "@openclaw Thread memory check: what is the hidden thread codename stored only in memory? Use memory tools first and reply only in this thread."
|
||||
promptSnippet: "Thread memory check"
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: answers the memory-backed fact inside the thread only
|
||||
actions:
|
||||
- call: reset
|
||||
- call: fs.writeFile
|
||||
args:
|
||||
- expr: "path.join(env.gateway.workspaceDir, 'MEMORY.md')"
|
||||
- expr: "`${config.memoryFact}\\n`"
|
||||
- utf8
|
||||
- call: forceMemoryIndex
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
query:
|
||||
expr: config.memoryQuery
|
||||
expectedNeedle:
|
||||
expr: config.expectedNeedle
|
||||
- call: handleQaAction
|
||||
saveAs: threadPayload
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
action: thread-create
|
||||
args:
|
||||
channelId:
|
||||
expr: config.channelId
|
||||
title:
|
||||
expr: config.threadTitle
|
||||
- set: threadId
|
||||
value:
|
||||
expr: "threadPayload?.thread?.id"
|
||||
- assert:
|
||||
expr: Boolean(threadId)
|
||||
message: missing thread id for memory isolation check
|
||||
- set: beforeCursor
|
||||
value:
|
||||
expr: state.getSnapshot().messages.length
|
||||
- call: state.addInboundMessage
|
||||
args:
|
||||
- conversation:
|
||||
id:
|
||||
expr: config.channelId
|
||||
kind: channel
|
||||
title:
|
||||
expr: config.channelTitle
|
||||
senderId: alice
|
||||
senderName: Alice
|
||||
text:
|
||||
expr: config.prompt
|
||||
threadId:
|
||||
ref: threadId
|
||||
threadTitle:
|
||||
expr: config.threadTitle
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: outbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === config.channelId && candidate.threadId === threadId && candidate.text.includes(config.expectedNeedle)"
|
||||
- expr: liveTurnTimeoutMs(env, 45000)
|
||||
- assert:
|
||||
expr: "!state.getSnapshot().messages.slice(beforeCursor).some((candidate) => candidate.direction === 'outbound' && candidate.conversation.id === config.channelId && !candidate.threadId)"
|
||||
message: threaded memory answer leaked into root channel
|
||||
- assert:
|
||||
expr: "!env.mock || (await fetchJson(`${env.mock.baseUrl}/debug/requests`)).filter((request) => String(request.allInputText ?? '').includes(config.promptSnippet)).some((request) => request.plannedToolName === 'memory_search')"
|
||||
message: expected memory_search in thread memory flow
|
||||
detailsExpr: outbound.text
|
||||
```
|
||||
90
openclaw/qa/scenarios/models/anthropic-opus-api-key-smoke.md
Normal file
90
openclaw/qa/scenarios/models/anthropic-opus-api-key-smoke.md
Normal file
|
|
@ -0,0 +1,90 @@
|
|||
# Anthropic Opus API key smoke
|
||||
|
||||
```yaml qa-scenario
|
||||
id: anthropic-opus-api-key-smoke
|
||||
title: Anthropic Opus API key smoke
|
||||
surface: model-provider
|
||||
coverage:
|
||||
primary:
|
||||
- models.provider-auth
|
||||
secondary:
|
||||
- models.anthropic
|
||||
objective: Verify the regular Anthropic Opus lane can complete a quick chat turn using API-key auth.
|
||||
successCriteria:
|
||||
- A live-frontier run fails fast unless the selected primary provider is anthropic.
|
||||
- The selected primary model is Anthropic Opus 4.6.
|
||||
- The QA gateway worker has an Anthropic API key available through environment auth.
|
||||
- The agent replies through the regular Anthropic provider.
|
||||
docsRefs:
|
||||
- docs/concepts/model-providers.md
|
||||
- docs/help/testing.md
|
||||
codeRefs:
|
||||
- extensions/anthropic/register.runtime.ts
|
||||
- extensions/qa-lab/src/gateway-child.ts
|
||||
- extensions/qa-lab/src/suite.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Run with `pnpm openclaw qa suite --provider-mode live-frontier --model anthropic/claude-opus-4-6 --alt-model anthropic/claude-opus-4-6 --scenario anthropic-opus-api-key-smoke`.
|
||||
config:
|
||||
requiredProvider: anthropic
|
||||
requiredModel: claude-opus-4-6
|
||||
chatPrompt: "Anthropic Opus API key smoke. Reply exactly: ANTHROPIC-OPUS-API-KEY-OK"
|
||||
chatExpected: ANTHROPIC-OPUS-API-KEY-OK
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: confirms regular Anthropic API-key lane
|
||||
actions:
|
||||
- set: selected
|
||||
value:
|
||||
expr: splitModelRef(env.primaryModel)
|
||||
- assert:
|
||||
expr: "env.providerMode !== 'live-frontier' || selected?.provider === config.requiredProvider"
|
||||
message:
|
||||
expr: "`expected live primary provider ${config.requiredProvider}, got ${env.primaryModel}`"
|
||||
- assert:
|
||||
expr: "env.providerMode !== 'live-frontier' || selected?.model === config.requiredModel"
|
||||
message:
|
||||
expr: "`expected live primary model ${config.requiredModel}, got ${env.primaryModel}`"
|
||||
- assert:
|
||||
expr: "env.providerMode !== 'live-frontier' || Boolean(env.gateway.runtimeEnv.ANTHROPIC_API_KEY?.trim())"
|
||||
message: expected ANTHROPIC_API_KEY to be available for API-key QA mode
|
||||
detailsExpr: "env.providerMode === 'live-frontier' ? `provider=${selected?.provider} model=${selected?.model} auth=env-api-key` : `mock-compatible provider=${selected?.provider}`"
|
||||
- name: talks through regular Anthropic Opus
|
||||
actions:
|
||||
- if:
|
||||
expr: "env.providerMode !== 'live-frontier'"
|
||||
then:
|
||||
- assert: "true"
|
||||
else:
|
||||
- call: reset
|
||||
- set: selected
|
||||
value:
|
||||
expr: splitModelRef(env.primaryModel)
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:anthropic-opus-api-key
|
||||
message:
|
||||
expr: config.chatPrompt
|
||||
provider:
|
||||
expr: selected?.provider
|
||||
model:
|
||||
expr: selected?.model
|
||||
timeoutMs:
|
||||
expr: resolveQaLiveTurnTimeoutMs(env, 60000, env.primaryModel)
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: chatOutbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === 'qa-operator'"
|
||||
- expr: resolveQaLiveTurnTimeoutMs(env, 30000, env.primaryModel)
|
||||
- assert:
|
||||
expr: "chatOutbound.text.includes(config.chatExpected)"
|
||||
message:
|
||||
expr: "`chat marker missing: ${chatOutbound.text}`"
|
||||
detailsExpr: "env.providerMode !== 'live-frontier' ? 'mock mode: skipped live Anthropic smoke' : chatOutbound.text"
|
||||
```
|
||||
|
|
@ -0,0 +1,95 @@
|
|||
# Anthropic Opus setup-token smoke
|
||||
|
||||
```yaml qa-scenario
|
||||
id: anthropic-opus-setup-token-smoke
|
||||
title: Anthropic Opus setup-token smoke
|
||||
surface: model-provider
|
||||
coverage:
|
||||
primary:
|
||||
- models.provider-auth
|
||||
secondary:
|
||||
- models.anthropic
|
||||
objective: Verify the regular Anthropic Opus lane can complete a quick chat turn using setup-token auth.
|
||||
successCriteria:
|
||||
- A live-frontier run fails fast unless the selected primary provider is anthropic.
|
||||
- The selected primary model is Anthropic Opus 4.6.
|
||||
- The QA gateway worker stages a token auth profile in the isolated agent store.
|
||||
- The agent replies through the regular Anthropic provider.
|
||||
docsRefs:
|
||||
- docs/concepts/model-providers.md
|
||||
- docs/help/testing.md
|
||||
codeRefs:
|
||||
- extensions/anthropic/register.runtime.ts
|
||||
- extensions/qa-lab/src/gateway-child.ts
|
||||
- extensions/qa-lab/src/suite.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Run with `OPENCLAW_LIVE_SETUP_TOKEN_VALUE=<setup-token> pnpm openclaw qa suite --provider-mode live-frontier --model anthropic/claude-opus-4-6 --alt-model anthropic/claude-opus-4-6 --scenario anthropic-opus-setup-token-smoke`.
|
||||
config:
|
||||
requiredProvider: anthropic
|
||||
requiredModel: claude-opus-4-6
|
||||
profileId: "anthropic:qa-setup-token"
|
||||
chatPrompt: "Anthropic Opus setup-token smoke. Reply exactly: ANTHROPIC-OPUS-SETUP-TOKEN-OK"
|
||||
chatExpected: ANTHROPIC-OPUS-SETUP-TOKEN-OK
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: confirms regular Anthropic setup-token lane
|
||||
actions:
|
||||
- set: selected
|
||||
value:
|
||||
expr: splitModelRef(env.primaryModel)
|
||||
- assert:
|
||||
expr: "env.providerMode !== 'live-frontier' || selected?.provider === config.requiredProvider"
|
||||
message:
|
||||
expr: "`expected live primary provider ${config.requiredProvider}, got ${env.primaryModel}`"
|
||||
- assert:
|
||||
expr: "env.providerMode !== 'live-frontier' || selected?.model === config.requiredModel"
|
||||
message:
|
||||
expr: "`expected live primary model ${config.requiredModel}, got ${env.primaryModel}`"
|
||||
- assert:
|
||||
expr: "env.providerMode !== 'live-frontier' || env.gateway.cfg.auth?.profiles?.[config.profileId]?.mode === 'token'"
|
||||
message:
|
||||
expr: "`expected token profile ${config.profileId} in QA config`"
|
||||
- assert:
|
||||
expr: "env.providerMode !== 'live-frontier' || !env.gateway.runtimeEnv.OPENCLAW_LIVE_SETUP_TOKEN_VALUE"
|
||||
message: setup-token value should not be passed to the gateway child env
|
||||
detailsExpr: "env.providerMode === 'live-frontier' ? `provider=${selected?.provider} model=${selected?.model} auth=setup-token profile=${config.profileId}` : `mock-compatible provider=${selected?.provider}`"
|
||||
- name: talks through regular Anthropic Opus
|
||||
actions:
|
||||
- if:
|
||||
expr: "env.providerMode !== 'live-frontier'"
|
||||
then:
|
||||
- assert: "true"
|
||||
else:
|
||||
- call: reset
|
||||
- set: selected
|
||||
value:
|
||||
expr: splitModelRef(env.primaryModel)
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:anthropic-opus-setup-token
|
||||
message:
|
||||
expr: config.chatPrompt
|
||||
provider:
|
||||
expr: selected?.provider
|
||||
model:
|
||||
expr: selected?.model
|
||||
timeoutMs:
|
||||
expr: resolveQaLiveTurnTimeoutMs(env, 60000, env.primaryModel)
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: chatOutbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === 'qa-operator'"
|
||||
- expr: resolveQaLiveTurnTimeoutMs(env, 30000, env.primaryModel)
|
||||
- assert:
|
||||
expr: "chatOutbound.text.includes(config.chatExpected)"
|
||||
message:
|
||||
expr: "`chat marker missing: ${chatOutbound.text}`"
|
||||
detailsExpr: "env.providerMode !== 'live-frontier' ? 'mock mode: skipped live Anthropic smoke' : chatOutbound.text"
|
||||
```
|
||||
|
|
@ -0,0 +1,258 @@
|
|||
# Claude CLI provider capabilities subscription
|
||||
|
||||
```yaml qa-scenario
|
||||
id: claude-cli-provider-capabilities-subscription
|
||||
title: Claude CLI provider capabilities subscription
|
||||
surface: model-provider
|
||||
coverage:
|
||||
primary:
|
||||
- models.provider-capabilities
|
||||
secondary:
|
||||
- models.claude-cli
|
||||
objective: Verify the Claude CLI model-provider lane can use native Claude subscription auth to talk, read an attached image, use bundled MCP tools, and apply workspace skills.
|
||||
successCriteria:
|
||||
- A live-frontier run fails fast unless the selected primary provider is claude-cli.
|
||||
- The Claude CLI backend does not preserve ANTHROPIC_API_KEY for this run, forcing native Claude subscription auth.
|
||||
- The agent replies through the Claude CLI provider in a direct chat turn.
|
||||
- The agent describes an attached image through the Claude CLI image path.
|
||||
- The agent can reach memory via the bundled MCP/tool bridge.
|
||||
- The agent sees and follows a workspace skill.
|
||||
docsRefs:
|
||||
- docs/gateway/cli-backends.md
|
||||
- docs/tools/skills.md
|
||||
- docs/cli/mcp.md
|
||||
- docs/tools/index.md
|
||||
codeRefs:
|
||||
- extensions/anthropic/cli-backend.ts
|
||||
- src/agents/cli-backends.ts
|
||||
- src/mcp/plugin-tools-serve.ts
|
||||
- extensions/qa-lab/src/suite.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Run with `pnpm openclaw qa suite --provider-mode live-frontier --cli-auth-mode subscription --model claude-cli/claude-sonnet-4-6 --alt-model claude-cli/claude-sonnet-4-6 --scenario claude-cli-provider-capabilities-subscription`.
|
||||
config:
|
||||
authMode: subscription
|
||||
requiredProvider: claude-cli
|
||||
chatPrompt: "Claude CLI provider marker check. Reply exactly: CLAUDE-CLI-CHAT-OK"
|
||||
chatExpected: CLAUDE-CLI-CHAT-OK
|
||||
imagePrompt: "Image understanding check: describe the top and bottom colors in the attached image in one short sentence."
|
||||
imageColorGroups:
|
||||
- [red, scarlet, crimson]
|
||||
- [blue, azure, teal, cyan, aqua]
|
||||
memoryFact: "Hidden Claude CLI MCP fact: the provider bridge codename is ORBIT-9."
|
||||
memoryQuery: "provider bridge codename ORBIT-9"
|
||||
memoryExpected: ORBIT-9
|
||||
memoryPrompt: "Memory tools check: use the available memory search MCP/tool bridge to find the hidden provider bridge codename stored only in memory. Reply with the codename."
|
||||
memoryPromptSnippet: "Memory tools check"
|
||||
skillName: qa-claude-cli-skill
|
||||
skillExpected: VISIBLE-SKILL-OK
|
||||
skillBody: |-
|
||||
---
|
||||
name: qa-claude-cli-skill
|
||||
description: Claude CLI QA skill marker
|
||||
---
|
||||
When the user asks for the Claude CLI skill marker exactly, or explicitly asks you to use qa-claude-cli-skill, reply with exactly: VISIBLE-SKILL-OK
|
||||
skillPrompt: "Use qa-claude-cli-skill now. Reply exactly with the visible skill marker and nothing else."
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: confirms the selected live provider and Claude CLI auth mode
|
||||
actions:
|
||||
- set: selected
|
||||
value:
|
||||
expr: splitModelRef(env.primaryModel)
|
||||
- set: preserveEnv
|
||||
value:
|
||||
expr: "String(env.gateway.runtimeEnv.OPENCLAW_LIVE_CLI_BACKEND_PRESERVE_ENV ?? '')"
|
||||
- assert:
|
||||
expr: "env.providerMode !== 'live-frontier' || selected?.provider === config.requiredProvider"
|
||||
message:
|
||||
expr: "`expected live primary provider ${config.requiredProvider}, got ${env.primaryModel}`"
|
||||
- assert:
|
||||
expr: "env.providerMode !== 'live-frontier' || env.gateway.runtimeEnv.OPENCLAW_LIVE_CLI_BACKEND_AUTH_MODE === config.authMode"
|
||||
message:
|
||||
expr: "`expected Claude CLI auth mode ${config.authMode}, got ${env.gateway.runtimeEnv.OPENCLAW_LIVE_CLI_BACKEND_AUTH_MODE ?? 'unset'}`"
|
||||
- assert:
|
||||
expr: "env.providerMode !== 'live-frontier' || !preserveEnv.includes('ANTHROPIC_API_KEY')"
|
||||
message:
|
||||
expr: "`expected ANTHROPIC_API_KEY not to be preserved for Claude CLI subscription QA mode, got ${preserveEnv}`"
|
||||
detailsExpr: "env.providerMode === 'live-frontier' ? `provider=${selected?.provider} model=${selected?.model} auth=${env.gateway.runtimeEnv.OPENCLAW_LIVE_CLI_BACKEND_AUTH_MODE} preserve=${preserveEnv}` : `mock-compatible provider=${selected?.provider}`"
|
||||
- name: talks through the selected provider
|
||||
actions:
|
||||
- call: reset
|
||||
- set: selected
|
||||
value:
|
||||
expr: splitModelRef(env.primaryModel)
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:claude-cli-chat
|
||||
message:
|
||||
expr: config.chatPrompt
|
||||
provider:
|
||||
expr: selected?.provider
|
||||
model:
|
||||
expr: selected?.model
|
||||
timeoutMs:
|
||||
expr: resolveQaLiveTurnTimeoutMs(env, 45000, env.primaryModel)
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: chatOutbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === 'qa-operator'"
|
||||
- expr: resolveQaLiveTurnTimeoutMs(env, 20000, env.primaryModel)
|
||||
- assert:
|
||||
expr: "chatOutbound.text.includes(config.chatExpected)"
|
||||
message:
|
||||
expr: "`chat marker missing: ${chatOutbound.text}`"
|
||||
detailsExpr: chatOutbound.text
|
||||
- name: describes an attached image through the selected provider
|
||||
actions:
|
||||
- call: reset
|
||||
- set: selected
|
||||
value:
|
||||
expr: splitModelRef(env.primaryModel)
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:claude-cli-image
|
||||
message:
|
||||
expr: config.imagePrompt
|
||||
provider:
|
||||
expr: selected?.provider
|
||||
model:
|
||||
expr: selected?.model
|
||||
attachments:
|
||||
- mimeType: image/png
|
||||
fileName: claude-cli-red-top-blue-bottom.png
|
||||
content:
|
||||
expr: imageUnderstandingValidPngBase64
|
||||
timeoutMs:
|
||||
expr: resolveQaLiveTurnTimeoutMs(env, 60000, env.primaryModel)
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: imageOutbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === 'qa-operator'"
|
||||
- expr: resolveQaLiveTurnTimeoutMs(env, 30000, env.primaryModel)
|
||||
- assert:
|
||||
expr: "config.imageColorGroups.every((group) => group.some((color) => normalizeLowercaseStringOrEmpty(imageOutbound.text).includes(color)))"
|
||||
message:
|
||||
expr: "`missing expected image colors: ${imageOutbound.text}`"
|
||||
- assert:
|
||||
expr: "!env.mock || (((await fetchJson(`${env.mock.baseUrl}/debug/requests`)).find((request) => String(request.prompt ?? '').includes('Image understanding check'))?.imageInputCount ?? 0) >= 1)"
|
||||
message: expected image input to reach mock provider
|
||||
detailsExpr: imageOutbound.text
|
||||
- name: reaches memory through the MCP/tool bridge
|
||||
actions:
|
||||
- call: reset
|
||||
- call: fs.writeFile
|
||||
args:
|
||||
- expr: "path.join(env.gateway.workspaceDir, 'MEMORY.md')"
|
||||
- expr: "`${config.memoryFact}\\n`"
|
||||
- utf8
|
||||
- call: forceMemoryIndex
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
query:
|
||||
expr: config.memoryQuery
|
||||
expectedNeedle:
|
||||
expr: config.memoryExpected
|
||||
- call: createSession
|
||||
saveAs: mcpSessionKey
|
||||
args:
|
||||
- ref: env
|
||||
- Claude CLI MCP bridge
|
||||
- call: readEffectiveTools
|
||||
saveAs: mcpTools
|
||||
args:
|
||||
- ref: env
|
||||
- ref: mcpSessionKey
|
||||
- assert:
|
||||
expr: "mcpTools.has('memory_search')"
|
||||
message: memory_search missing from effective tools before MCP bridge check
|
||||
- set: selected
|
||||
value:
|
||||
expr: splitModelRef(env.primaryModel)
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey:
|
||||
ref: mcpSessionKey
|
||||
message:
|
||||
expr: config.memoryPrompt
|
||||
provider:
|
||||
expr: selected?.provider
|
||||
model:
|
||||
expr: selected?.model
|
||||
timeoutMs:
|
||||
expr: resolveQaLiveTurnTimeoutMs(env, 90000, env.primaryModel)
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: mcpOutbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === 'qa-operator'"
|
||||
- expr: resolveQaLiveTurnTimeoutMs(env, 45000, env.primaryModel)
|
||||
- assert:
|
||||
expr: "mcpOutbound.text.includes(config.memoryExpected)"
|
||||
message:
|
||||
expr: "`MCP memory result missing ${config.memoryExpected}: ${mcpOutbound.text}`"
|
||||
- assert:
|
||||
expr: "!env.mock || (await fetchJson(`${env.mock.baseUrl}/debug/requests`)).filter((request) => String(request.allInputText ?? '').includes(config.memoryPromptSnippet)).some((request) => request.plannedToolName === 'memory_search')"
|
||||
message: expected mock model to plan memory_search for MCP bridge prompt
|
||||
detailsExpr: mcpOutbound.text
|
||||
- name: applies a workspace skill through the selected provider
|
||||
actions:
|
||||
- call: reset
|
||||
- call: writeWorkspaceSkill
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
name:
|
||||
expr: config.skillName
|
||||
body:
|
||||
expr: config.skillBody
|
||||
- call: waitForCondition
|
||||
args:
|
||||
- lambda:
|
||||
async: true
|
||||
expr: "((await readSkillStatus(env)).find((skill) => skill.name === config.skillName)?.eligible ? true : undefined)"
|
||||
- 15000
|
||||
- 200
|
||||
- set: selected
|
||||
value:
|
||||
expr: splitModelRef(env.primaryModel)
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:claude-cli-skill
|
||||
message:
|
||||
expr: config.skillPrompt
|
||||
provider:
|
||||
expr: selected?.provider
|
||||
model:
|
||||
expr: selected?.model
|
||||
timeoutMs:
|
||||
expr: resolveQaLiveTurnTimeoutMs(env, 60000, env.primaryModel)
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: skillOutbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === 'qa-operator'"
|
||||
- expr: resolveQaLiveTurnTimeoutMs(env, 30000, env.primaryModel)
|
||||
- assert:
|
||||
expr: "skillOutbound.text.includes(config.skillExpected)"
|
||||
message:
|
||||
expr: "`skill marker missing: ${skillOutbound.text}`"
|
||||
detailsExpr: skillOutbound.text
|
||||
```
|
||||
258
openclaw/qa/scenarios/models/claude-cli-provider-capabilities.md
Normal file
258
openclaw/qa/scenarios/models/claude-cli-provider-capabilities.md
Normal file
|
|
@ -0,0 +1,258 @@
|
|||
# Claude CLI provider capabilities API key
|
||||
|
||||
```yaml qa-scenario
|
||||
id: claude-cli-provider-capabilities
|
||||
title: Claude CLI provider capabilities API key
|
||||
surface: model-provider
|
||||
coverage:
|
||||
primary:
|
||||
- models.provider-capabilities
|
||||
secondary:
|
||||
- models.claude-cli
|
||||
objective: Verify the Claude CLI model-provider lane can use the Anthropic API key path to talk, read an attached image, use bundled MCP tools, and apply workspace skills.
|
||||
successCriteria:
|
||||
- A live-frontier run fails fast unless the selected primary provider is claude-cli.
|
||||
- The Claude CLI backend preserves ANTHROPIC_API_KEY for this run instead of using native subscription auth.
|
||||
- The agent replies through the Claude CLI provider in a direct chat turn.
|
||||
- The agent describes an attached image through the Claude CLI image path.
|
||||
- The agent can reach memory via the bundled MCP/tool bridge.
|
||||
- The agent sees and follows a workspace skill.
|
||||
docsRefs:
|
||||
- docs/gateway/cli-backends.md
|
||||
- docs/tools/skills.md
|
||||
- docs/cli/mcp.md
|
||||
- docs/tools/index.md
|
||||
codeRefs:
|
||||
- extensions/anthropic/cli-backend.ts
|
||||
- src/agents/cli-backends.ts
|
||||
- src/mcp/plugin-tools-serve.ts
|
||||
- extensions/qa-lab/src/suite.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Run with `pnpm openclaw qa suite --provider-mode live-frontier --cli-auth-mode api-key --model claude-cli/claude-sonnet-4-6 --alt-model claude-cli/claude-sonnet-4-6 --scenario claude-cli-provider-capabilities`.
|
||||
config:
|
||||
authMode: api-key
|
||||
requiredProvider: claude-cli
|
||||
chatPrompt: "Claude CLI provider marker check. Reply exactly: CLAUDE-CLI-CHAT-OK"
|
||||
chatExpected: CLAUDE-CLI-CHAT-OK
|
||||
imagePrompt: "Image understanding check: describe the top and bottom colors in the attached image in one short sentence."
|
||||
imageColorGroups:
|
||||
- [red, scarlet, crimson]
|
||||
- [blue, azure, teal, cyan, aqua]
|
||||
memoryFact: "Hidden Claude CLI MCP fact: the provider bridge codename is ORBIT-9."
|
||||
memoryQuery: "provider bridge codename ORBIT-9"
|
||||
memoryExpected: ORBIT-9
|
||||
memoryPrompt: "Memory tools check: use the available memory search MCP/tool bridge to find the hidden provider bridge codename stored only in memory. Reply with the codename."
|
||||
memoryPromptSnippet: "Memory tools check"
|
||||
skillName: qa-claude-cli-skill
|
||||
skillExpected: VISIBLE-SKILL-OK
|
||||
skillBody: |-
|
||||
---
|
||||
name: qa-claude-cli-skill
|
||||
description: Claude CLI QA skill marker
|
||||
---
|
||||
When the user asks for the Claude CLI skill marker exactly, or explicitly asks you to use qa-claude-cli-skill, reply with exactly: VISIBLE-SKILL-OK
|
||||
skillPrompt: "Use qa-claude-cli-skill now. Reply exactly with the visible skill marker and nothing else."
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: confirms the selected live provider and Claude CLI auth mode
|
||||
actions:
|
||||
- set: selected
|
||||
value:
|
||||
expr: splitModelRef(env.primaryModel)
|
||||
- set: preserveEnv
|
||||
value:
|
||||
expr: "String(env.gateway.runtimeEnv.OPENCLAW_LIVE_CLI_BACKEND_PRESERVE_ENV ?? '')"
|
||||
- assert:
|
||||
expr: "env.providerMode !== 'live-frontier' || selected?.provider === config.requiredProvider"
|
||||
message:
|
||||
expr: "`expected live primary provider ${config.requiredProvider}, got ${env.primaryModel}`"
|
||||
- assert:
|
||||
expr: "env.providerMode !== 'live-frontier' || env.gateway.runtimeEnv.OPENCLAW_LIVE_CLI_BACKEND_AUTH_MODE === config.authMode"
|
||||
message:
|
||||
expr: "`expected Claude CLI auth mode ${config.authMode}, got ${env.gateway.runtimeEnv.OPENCLAW_LIVE_CLI_BACKEND_AUTH_MODE ?? 'unset'}`"
|
||||
- assert:
|
||||
expr: "env.providerMode !== 'live-frontier' || preserveEnv.includes('ANTHROPIC_API_KEY')"
|
||||
message:
|
||||
expr: "`expected ANTHROPIC_API_KEY to be preserved for Claude CLI API-key QA mode, got ${preserveEnv}`"
|
||||
detailsExpr: "env.providerMode === 'live-frontier' ? `provider=${selected?.provider} model=${selected?.model} auth=${env.gateway.runtimeEnv.OPENCLAW_LIVE_CLI_BACKEND_AUTH_MODE} preserve=${preserveEnv}` : `mock-compatible provider=${selected?.provider}`"
|
||||
- name: talks through the selected provider
|
||||
actions:
|
||||
- call: reset
|
||||
- set: selected
|
||||
value:
|
||||
expr: splitModelRef(env.primaryModel)
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:claude-cli-chat
|
||||
message:
|
||||
expr: config.chatPrompt
|
||||
provider:
|
||||
expr: selected?.provider
|
||||
model:
|
||||
expr: selected?.model
|
||||
timeoutMs:
|
||||
expr: resolveQaLiveTurnTimeoutMs(env, 45000, env.primaryModel)
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: chatOutbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === 'qa-operator'"
|
||||
- expr: resolveQaLiveTurnTimeoutMs(env, 20000, env.primaryModel)
|
||||
- assert:
|
||||
expr: "chatOutbound.text.includes(config.chatExpected)"
|
||||
message:
|
||||
expr: "`chat marker missing: ${chatOutbound.text}`"
|
||||
detailsExpr: chatOutbound.text
|
||||
- name: describes an attached image through the selected provider
|
||||
actions:
|
||||
- call: reset
|
||||
- set: selected
|
||||
value:
|
||||
expr: splitModelRef(env.primaryModel)
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:claude-cli-image
|
||||
message:
|
||||
expr: config.imagePrompt
|
||||
provider:
|
||||
expr: selected?.provider
|
||||
model:
|
||||
expr: selected?.model
|
||||
attachments:
|
||||
- mimeType: image/png
|
||||
fileName: claude-cli-red-top-blue-bottom.png
|
||||
content:
|
||||
expr: imageUnderstandingValidPngBase64
|
||||
timeoutMs:
|
||||
expr: resolveQaLiveTurnTimeoutMs(env, 60000, env.primaryModel)
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: imageOutbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === 'qa-operator'"
|
||||
- expr: resolveQaLiveTurnTimeoutMs(env, 30000, env.primaryModel)
|
||||
- assert:
|
||||
expr: "config.imageColorGroups.every((group) => group.some((color) => normalizeLowercaseStringOrEmpty(imageOutbound.text).includes(color)))"
|
||||
message:
|
||||
expr: "`missing expected image colors: ${imageOutbound.text}`"
|
||||
- assert:
|
||||
expr: "!env.mock || (((await fetchJson(`${env.mock.baseUrl}/debug/requests`)).find((request) => String(request.prompt ?? '').includes('Image understanding check'))?.imageInputCount ?? 0) >= 1)"
|
||||
message: expected image input to reach mock provider
|
||||
detailsExpr: imageOutbound.text
|
||||
- name: reaches memory through the MCP/tool bridge
|
||||
actions:
|
||||
- call: reset
|
||||
- call: fs.writeFile
|
||||
args:
|
||||
- expr: "path.join(env.gateway.workspaceDir, 'MEMORY.md')"
|
||||
- expr: "`${config.memoryFact}\\n`"
|
||||
- utf8
|
||||
- call: forceMemoryIndex
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
query:
|
||||
expr: config.memoryQuery
|
||||
expectedNeedle:
|
||||
expr: config.memoryExpected
|
||||
- call: createSession
|
||||
saveAs: mcpSessionKey
|
||||
args:
|
||||
- ref: env
|
||||
- Claude CLI MCP bridge
|
||||
- call: readEffectiveTools
|
||||
saveAs: mcpTools
|
||||
args:
|
||||
- ref: env
|
||||
- ref: mcpSessionKey
|
||||
- assert:
|
||||
expr: "mcpTools.has('memory_search')"
|
||||
message: memory_search missing from effective tools before MCP bridge check
|
||||
- set: selected
|
||||
value:
|
||||
expr: splitModelRef(env.primaryModel)
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey:
|
||||
ref: mcpSessionKey
|
||||
message:
|
||||
expr: config.memoryPrompt
|
||||
provider:
|
||||
expr: selected?.provider
|
||||
model:
|
||||
expr: selected?.model
|
||||
timeoutMs:
|
||||
expr: resolveQaLiveTurnTimeoutMs(env, 90000, env.primaryModel)
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: mcpOutbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === 'qa-operator'"
|
||||
- expr: resolveQaLiveTurnTimeoutMs(env, 45000, env.primaryModel)
|
||||
- assert:
|
||||
expr: "mcpOutbound.text.includes(config.memoryExpected)"
|
||||
message:
|
||||
expr: "`MCP memory result missing ${config.memoryExpected}: ${mcpOutbound.text}`"
|
||||
- assert:
|
||||
expr: "!env.mock || (await fetchJson(`${env.mock.baseUrl}/debug/requests`)).filter((request) => String(request.allInputText ?? '').includes(config.memoryPromptSnippet)).some((request) => request.plannedToolName === 'memory_search')"
|
||||
message: expected mock model to plan memory_search for MCP bridge prompt
|
||||
detailsExpr: mcpOutbound.text
|
||||
- name: applies a workspace skill through the selected provider
|
||||
actions:
|
||||
- call: reset
|
||||
- call: writeWorkspaceSkill
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
name:
|
||||
expr: config.skillName
|
||||
body:
|
||||
expr: config.skillBody
|
||||
- call: waitForCondition
|
||||
args:
|
||||
- lambda:
|
||||
async: true
|
||||
expr: "((await readSkillStatus(env)).find((skill) => skill.name === config.skillName)?.eligible ? true : undefined)"
|
||||
- 15000
|
||||
- 200
|
||||
- set: selected
|
||||
value:
|
||||
expr: splitModelRef(env.primaryModel)
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:claude-cli-skill
|
||||
message:
|
||||
expr: config.skillPrompt
|
||||
provider:
|
||||
expr: selected?.provider
|
||||
model:
|
||||
expr: selected?.model
|
||||
timeoutMs:
|
||||
expr: resolveQaLiveTurnTimeoutMs(env, 60000, env.primaryModel)
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: skillOutbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === 'qa-operator'"
|
||||
- expr: resolveQaLiveTurnTimeoutMs(env, 30000, env.primaryModel)
|
||||
- assert:
|
||||
expr: "skillOutbound.text.includes(config.skillExpected)"
|
||||
message:
|
||||
expr: "`skill marker missing: ${skillOutbound.text}`"
|
||||
detailsExpr: skillOutbound.text
|
||||
```
|
||||
147
openclaw/qa/scenarios/models/codex-harness-no-meta-leak.md
Normal file
147
openclaw/qa/scenarios/models/codex-harness-no-meta-leak.md
Normal file
|
|
@ -0,0 +1,147 @@
|
|||
# Codex harness no meta leak
|
||||
|
||||
```yaml qa-scenario
|
||||
id: codex-harness-no-meta-leak
|
||||
title: Codex harness no meta leak
|
||||
surface: dm
|
||||
coverage:
|
||||
primary:
|
||||
- models.codex-cli
|
||||
secondary:
|
||||
- runtime.no-meta-leak
|
||||
objective: Verify the Codex app-server harness keeps coordination/meta chatter out of the visible reply.
|
||||
successCriteria:
|
||||
- The scenario forces the Codex embedded harness and disables PI fallback.
|
||||
- The final visible reply includes the requested confirmation token.
|
||||
- The visible reply does not include internal coordination or progress chatter.
|
||||
docsRefs:
|
||||
- docs/plugins/sdk-agent-harness.md
|
||||
- docs/concepts/qa-e2e-automation.md
|
||||
- docs/help/testing.md
|
||||
codeRefs:
|
||||
- extensions/codex/src/app-server/event-projector.ts
|
||||
- extensions/qa-lab/src/reply-failure.ts
|
||||
- extensions/qa-lab/src/suite.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Run with `pnpm openclaw qa suite --provider-mode live-frontier --model codex/gpt-5.4 --alt-model codex/gpt-5.4 --scenario codex-harness-no-meta-leak`.
|
||||
config:
|
||||
requiredProvider: codex
|
||||
requiredModel: gpt-5.4
|
||||
harnessRuntime: codex
|
||||
harnessFallback: none
|
||||
expectedReply: QA_LEAK_OK
|
||||
prompt: |-
|
||||
Think through your answer privately, but do not expose any internal planning, thread-context checks, or progress narration.
|
||||
|
||||
Reply with exactly one line:
|
||||
QA_LEAK_OK
|
||||
forbiddenReplySubstrings:
|
||||
- checking thread context
|
||||
- thread context thin
|
||||
- post a tight progress reply here
|
||||
- posting a coordination nudge
|
||||
- posted a short coordination reply
|
||||
- not inventing status
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: confirms GPT-5.4 Codex harness target
|
||||
actions:
|
||||
- set: selected
|
||||
value:
|
||||
expr: splitModelRef(env.primaryModel)
|
||||
- assert:
|
||||
expr: "env.providerMode !== 'live-frontier' || selected?.provider === config.requiredProvider"
|
||||
message:
|
||||
expr: "`expected live primary provider ${config.requiredProvider}, got ${env.primaryModel}`"
|
||||
- assert:
|
||||
expr: "env.providerMode !== 'live-frontier' || selected?.model === config.requiredModel"
|
||||
message:
|
||||
expr: "`expected live primary model ${config.requiredModel}, got ${env.primaryModel}`"
|
||||
- if:
|
||||
expr: "env.providerMode !== 'live-frontier'"
|
||||
then:
|
||||
- assert: "true"
|
||||
else:
|
||||
- call: patchConfig
|
||||
saveAs: patchResult
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
patch:
|
||||
agents:
|
||||
defaults:
|
||||
embeddedHarness:
|
||||
runtime:
|
||||
expr: config.harnessRuntime
|
||||
fallback:
|
||||
expr: config.harnessFallback
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- call: waitForQaChannelReady
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- call: readConfigSnapshot
|
||||
saveAs: snapshot
|
||||
args:
|
||||
- ref: env
|
||||
- assert:
|
||||
expr: "snapshot.config.agents?.defaults?.embeddedHarness?.runtime === config.harnessRuntime"
|
||||
message:
|
||||
expr: "`expected embeddedHarness.runtime=${config.harnessRuntime}, got ${JSON.stringify(snapshot.config.agents?.defaults?.embeddedHarness)}`"
|
||||
- assert:
|
||||
expr: "snapshot.config.agents?.defaults?.embeddedHarness?.fallback === config.harnessFallback"
|
||||
message:
|
||||
expr: "`expected embeddedHarness.fallback=${config.harnessFallback}, got ${JSON.stringify(snapshot.config.agents?.defaults?.embeddedHarness)}`"
|
||||
detailsExpr: "env.providerMode === 'live-frontier' ? `provider=${selected?.provider} model=${selected?.model} runtime=${snapshot.config.agents?.defaults?.embeddedHarness?.runtime} fallback=${snapshot.config.agents?.defaults?.embeddedHarness?.fallback}` : `mock mode: parsed ${scenario.id}`"
|
||||
- name: keeps codex coordination chatter out of the visible reply
|
||||
actions:
|
||||
- if:
|
||||
expr: "env.providerMode !== 'live-frontier'"
|
||||
then:
|
||||
- assert: "true"
|
||||
else:
|
||||
- call: reset
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:codex-meta-leak
|
||||
message:
|
||||
expr: config.prompt
|
||||
provider:
|
||||
expr: selected?.provider
|
||||
model:
|
||||
expr: selected?.model
|
||||
timeoutMs:
|
||||
expr: resolveQaLiveTurnTimeoutMs(env, 180000, env.primaryModel)
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: outbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === 'qa-operator' && candidate.text.includes(config.expectedReply)"
|
||||
- expr: resolveQaLiveTurnTimeoutMs(env, 60000, env.primaryModel)
|
||||
- set: outboundLower
|
||||
value:
|
||||
expr: normalizeLowercaseStringOrEmpty(outbound.text)
|
||||
- assert:
|
||||
expr: "outbound.text.trim() === config.expectedReply"
|
||||
message:
|
||||
expr: "`expected exact visible reply ${config.expectedReply}, got ${outbound.text}`"
|
||||
- forEach:
|
||||
items:
|
||||
expr: "config.forbiddenReplySubstrings ?? []"
|
||||
item: forbidden
|
||||
actions:
|
||||
- assert:
|
||||
expr: "!outboundLower.includes(normalizeLowercaseStringOrEmpty(forbidden))"
|
||||
message:
|
||||
expr: "`visible reply leaked internal meta text (${forbidden}): ${outbound.text}`"
|
||||
detailsExpr: "env.providerMode !== 'live-frontier' ? 'mock mode: skipped live codex leak check' : outbound.text"
|
||||
```
|
||||
79
openclaw/qa/scenarios/models/model-switch-follow-up.md
Normal file
79
openclaw/qa/scenarios/models/model-switch-follow-up.md
Normal file
|
|
@ -0,0 +1,79 @@
|
|||
# Model switch follow-up
|
||||
|
||||
```yaml qa-scenario
|
||||
id: model-switch-follow-up
|
||||
title: Model switch follow-up
|
||||
surface: models
|
||||
coverage:
|
||||
primary:
|
||||
- models.switching
|
||||
secondary:
|
||||
- runtime.session-continuity
|
||||
objective: Verify the agent can switch to a different configured model and continue coherently.
|
||||
successCriteria:
|
||||
- Agent reflects the model switch request.
|
||||
- Follow-up answer remains coherent with prior context.
|
||||
- Final report notes whether the switch actually happened.
|
||||
docsRefs:
|
||||
- docs/help/testing.md
|
||||
- docs/web/dashboard.md
|
||||
codeRefs:
|
||||
- extensions/qa-lab/src/report.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify the agent can switch to a different configured model and continue coherently.
|
||||
config:
|
||||
initialPrompt: "Say hello from the default configured model."
|
||||
followupPrompt: "Continue the exchange after switching models and note the handoff."
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: runs on the default configured model
|
||||
actions:
|
||||
- call: reset
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:model-switch
|
||||
message:
|
||||
expr: config.initialPrompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 30000)
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: outbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === 'qa-operator'"
|
||||
detailsExpr: "env.mock ? String((await fetchJson(`${env.mock.baseUrl}/debug/last-request`))?.body?.model ?? '') : outbound.text"
|
||||
- name: switches to the alternate model and continues
|
||||
actions:
|
||||
- set: alternate
|
||||
value:
|
||||
expr: splitModelRef(env.alternateModel)
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:model-switch
|
||||
message:
|
||||
expr: config.followupPrompt
|
||||
provider:
|
||||
expr: alternate?.provider
|
||||
model:
|
||||
expr: alternate?.model
|
||||
timeoutMs:
|
||||
expr: resolveQaLiveTurnTimeoutMs(env, 30000, env.alternateModel)
|
||||
- call: waitForCondition
|
||||
saveAs: outbound
|
||||
args:
|
||||
- lambda:
|
||||
expr: "state.getSnapshot().messages.filter((candidate) => candidate.direction === 'outbound' && candidate.conversation.id === 'qa-operator' && (() => { const lower = normalizeLowercaseStringOrEmpty(candidate.text); return lower.includes('switch') || lower.includes('handoff'); })()).at(-1)"
|
||||
- expr: resolveQaLiveTurnTimeoutMs(env, 20000, env.alternateModel)
|
||||
- assert:
|
||||
expr: "!env.mock || ((await fetchJson(`${env.mock.baseUrl}/debug/last-request`))?.body?.model === 'gpt-5.4-alt')"
|
||||
message:
|
||||
expr: "`expected gpt-5.4-alt, got ${String((await fetchJson(`${env.mock.baseUrl}/debug/last-request`))?.body?.model ?? '')}`"
|
||||
detailsExpr: outbound.text
|
||||
```
|
||||
95
openclaw/qa/scenarios/models/model-switch-tool-continuity.md
Normal file
95
openclaw/qa/scenarios/models/model-switch-tool-continuity.md
Normal file
|
|
@ -0,0 +1,95 @@
|
|||
# Model switch with tool continuity
|
||||
|
||||
```yaml qa-scenario
|
||||
id: model-switch-tool-continuity
|
||||
title: Model switch with tool continuity
|
||||
surface: models
|
||||
coverage:
|
||||
primary:
|
||||
- models.switching
|
||||
secondary:
|
||||
- runtime.tool-continuity
|
||||
objective: Verify switching models preserves session context and tool use instead of dropping into plain-text only behavior.
|
||||
successCriteria:
|
||||
- Alternate model is actually requested.
|
||||
- A tool call still happens after the model switch.
|
||||
- Final answer acknowledges the handoff and uses the tool-derived evidence.
|
||||
docsRefs:
|
||||
- docs/help/testing.md
|
||||
- docs/concepts/model-failover.md
|
||||
codeRefs:
|
||||
- extensions/qa-lab/src/suite.ts
|
||||
- extensions/qa-lab/src/mock-openai-server.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify switching models preserves session context and tool use instead of dropping into plain-text only behavior.
|
||||
config:
|
||||
initialPrompt: "Read repo/qa/scenarios/index.md and summarize the QA scenario pack mission in one clause before any model switch."
|
||||
followupPrompt: "The harness has already requested the alternate model for this turn. Do not call session_status or change models yourself. Tool continuity check: use the read tool to reread repo/qa/scenarios/index.md, then mention the model handoff and QA mission in one short sentence."
|
||||
promptSnippet: "Tool continuity check"
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: keeps using tools after switching models
|
||||
actions:
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- call: reset
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:model-switch-tools
|
||||
message:
|
||||
expr: config.initialPrompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 30000)
|
||||
- set: alternate
|
||||
value:
|
||||
expr: splitModelRef(env.alternateModel)
|
||||
- set: beforeSwitchCursor
|
||||
value:
|
||||
expr: state.getSnapshot().messages.length
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:model-switch-tools
|
||||
message:
|
||||
expr: config.followupPrompt
|
||||
provider:
|
||||
expr: alternate?.provider
|
||||
model:
|
||||
expr: alternate?.model
|
||||
timeoutMs:
|
||||
expr: resolveQaLiveTurnTimeoutMs(env, 30000, env.alternateModel)
|
||||
- call: waitForCondition
|
||||
saveAs: outbound
|
||||
args:
|
||||
- lambda:
|
||||
expr: "state.getSnapshot().messages.slice(beforeSwitchCursor).filter((candidate) => candidate.direction === 'outbound' && candidate.conversation.id === 'qa-operator' && hasModelSwitchContinuityEvidence(candidate.text)).at(-1)"
|
||||
- expr: resolveQaLiveTurnTimeoutMs(env, 20000, env.alternateModel)
|
||||
- assert:
|
||||
expr: hasModelSwitchContinuityEvidence(outbound.text)
|
||||
message:
|
||||
expr: "`switch reply missed kickoff continuity: ${outbound.text}`"
|
||||
- if:
|
||||
expr: "Boolean(env.mock)"
|
||||
then:
|
||||
- set: switchDebugRequests
|
||||
value:
|
||||
expr: "await fetchJson(`${env.mock.baseUrl}/debug/requests`)"
|
||||
- set: switchRequest
|
||||
value:
|
||||
expr: "switchDebugRequests.find((request) => String(request.allInputText ?? '').includes(config.promptSnippet))"
|
||||
- assert:
|
||||
expr: "switchRequest?.plannedToolName === 'read'"
|
||||
message:
|
||||
expr: "`expected read after switch, got ${String(switchRequest?.plannedToolName ?? '')}`"
|
||||
- assert:
|
||||
expr: "String(switchRequest?.model ?? '') === String(alternate?.model ?? '')"
|
||||
message:
|
||||
expr: "`expected alternate model, got ${String(switchRequest?.model ?? '')}`"
|
||||
detailsExpr: outbound.text
|
||||
```
|
||||
126
openclaw/qa/scenarios/plugins/bundled-plugin-skill-runtime.md
Normal file
126
openclaw/qa/scenarios/plugins/bundled-plugin-skill-runtime.md
Normal file
|
|
@ -0,0 +1,126 @@
|
|||
# Bundled plugin skill runtime
|
||||
|
||||
```yaml qa-scenario
|
||||
id: bundled-plugin-skill-runtime
|
||||
title: Bundled plugin skill runtime
|
||||
surface: skills
|
||||
coverage:
|
||||
primary:
|
||||
- plugins.skills
|
||||
secondary:
|
||||
- plugins.runtime
|
||||
objective: Verify packaged bundled plugin skills load from dist-runtime instead of being skipped by path-containment checks.
|
||||
successCriteria:
|
||||
- The runtime-packaged bundled plugin tree is used as OPENCLAW_BUNDLED_PLUGINS_DIR.
|
||||
- The enabled bundled plugin skill is reported as eligible by the skills CLI.
|
||||
- The check fails on SKILL.md symlink escapes and passes when runtime staging copies SKILL.md as a real file.
|
||||
docsRefs:
|
||||
- docs/tools/skills.md
|
||||
- docs/plugins/manifest.md
|
||||
codeRefs:
|
||||
- scripts/stage-bundled-plugin-runtime.mjs
|
||||
- src/agents/skills/workspace.ts
|
||||
- src/agents/skills/plugin-skills.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Force the packaged dist-runtime plugin tree and verify an enabled bundled plugin skill survives discovery.
|
||||
config:
|
||||
pluginId: open-prose
|
||||
expectedSkillName: prose
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: loads a bundled plugin skill from dist-runtime
|
||||
actions:
|
||||
- set: skillCheck
|
||||
value:
|
||||
expr: |-
|
||||
(async () => {
|
||||
const { spawnSync } = await qaImport("node:child_process");
|
||||
const fsSync = await qaImport("node:fs");
|
||||
const distRuntimeExtensions = path.join(env.repoRoot, "dist-runtime", "extensions");
|
||||
const skillPath = path.join(
|
||||
distRuntimeExtensions,
|
||||
config.pluginId,
|
||||
"skills",
|
||||
config.expectedSkillName,
|
||||
"SKILL.md",
|
||||
);
|
||||
const tempRoot = await fs.mkdtemp(path.join(env.gateway.tempRoot, "bundled-skill-runtime-"));
|
||||
const homeDir = path.join(tempRoot, "home");
|
||||
const stateDir = path.join(tempRoot, "state");
|
||||
const workspaceDir = path.join(tempRoot, "workspace");
|
||||
const xdgConfigHome = path.join(tempRoot, "xdg-config");
|
||||
const xdgDataHome = path.join(tempRoot, "xdg-data");
|
||||
const xdgCacheHome = path.join(tempRoot, "xdg-cache");
|
||||
await Promise.all(
|
||||
[homeDir, stateDir, workspaceDir, xdgConfigHome, xdgDataHome, xdgCacheHome].map((dir) =>
|
||||
fs.mkdir(dir, { recursive: true }),
|
||||
),
|
||||
);
|
||||
const configPath = path.join(tempRoot, "openclaw.json");
|
||||
await fs.writeFile(
|
||||
configPath,
|
||||
`${JSON.stringify(
|
||||
{
|
||||
agents: { defaults: { workspace: workspaceDir } },
|
||||
plugins: {
|
||||
allow: [config.pluginId],
|
||||
entries: { [config.pluginId]: { enabled: true } },
|
||||
},
|
||||
},
|
||||
null,
|
||||
2,
|
||||
)}\n`,
|
||||
"utf8",
|
||||
);
|
||||
const cliEnv = {
|
||||
...env.gateway.runtimeEnv,
|
||||
HOME: homeDir,
|
||||
OPENCLAW_HOME: homeDir,
|
||||
OPENCLAW_CONFIG_PATH: configPath,
|
||||
OPENCLAW_STATE_DIR: stateDir,
|
||||
OPENCLAW_OAUTH_DIR: path.join(stateDir, "credentials"),
|
||||
OPENCLAW_BUNDLED_PLUGINS_DIR: distRuntimeExtensions,
|
||||
XDG_CONFIG_HOME: xdgConfigHome,
|
||||
XDG_DATA_HOME: xdgDataHome,
|
||||
XDG_CACHE_HOME: xdgCacheHome,
|
||||
};
|
||||
const result = spawnSync(
|
||||
process.execPath,
|
||||
[path.join(env.repoRoot, "dist", "index.js"), "skills", "list", "--json", "--eligible"],
|
||||
{
|
||||
cwd: tempRoot,
|
||||
env: cliEnv,
|
||||
encoding: "utf8",
|
||||
timeout: 60000,
|
||||
},
|
||||
);
|
||||
let parsed = null;
|
||||
let parseError = null;
|
||||
try {
|
||||
parsed = result.stdout ? JSON.parse(result.stdout) : null;
|
||||
} catch (error) {
|
||||
parseError = formatErrorMessage(error);
|
||||
}
|
||||
const skills = Array.isArray(parsed?.skills) ? parsed.skills : [];
|
||||
const skill = skills.find((entry) => entry?.name === config.expectedSkillName);
|
||||
return {
|
||||
exitCode: result.status,
|
||||
signal: result.signal,
|
||||
parseError,
|
||||
skill,
|
||||
skillNames: skills.map((entry) => entry?.name).filter(Boolean).sort(),
|
||||
skillPath: path.relative(env.repoRoot, skillPath),
|
||||
skillMdSymlink: fsSync.existsSync(skillPath) ? fsSync.lstatSync(skillPath).isSymbolicLink() : null,
|
||||
stderr: String(result.stderr ?? "").replaceAll(env.repoRoot, "<repo>").trim().slice(0, 1200),
|
||||
};
|
||||
})()
|
||||
- assert:
|
||||
expr: "skillCheck.exitCode === 0 && skillCheck.skill?.eligible === true && !skillCheck.skill?.disabled && !skillCheck.skill?.blockedByAllowlist"
|
||||
message:
|
||||
expr: |-
|
||||
`expected bundled plugin skill "${config.expectedSkillName}" from "${config.pluginId}" to load from dist-runtime; got ${JSON.stringify(skillCheck.skill)}; SKILL.md symlink=${skillCheck.skillMdSymlink}; stderr=${skillCheck.stderr || "(empty)"}`
|
||||
detailsExpr: skillCheck
|
||||
```
|
||||
67
openclaw/qa/scenarios/plugins/mcp-plugin-tools-call.md
Normal file
67
openclaw/qa/scenarios/plugins/mcp-plugin-tools-call.md
Normal file
|
|
@ -0,0 +1,67 @@
|
|||
# MCP plugin-tools call
|
||||
|
||||
```yaml qa-scenario
|
||||
id: mcp-plugin-tools-call
|
||||
title: MCP plugin-tools call
|
||||
surface: mcp
|
||||
coverage:
|
||||
primary:
|
||||
- plugins.mcp-tools
|
||||
secondary:
|
||||
- tools.invocation
|
||||
objective: Verify OpenClaw can expose plugin tools over MCP and a real MCP client can call one successfully.
|
||||
successCriteria:
|
||||
- Plugin tools MCP server lists memory_search.
|
||||
- A real MCP client calls memory_search successfully.
|
||||
- The returned MCP payload includes the expected memory-only fact.
|
||||
docsRefs:
|
||||
- docs/cli/mcp.md
|
||||
- docs/gateway/protocol.md
|
||||
codeRefs:
|
||||
- src/mcp/plugin-tools-serve.ts
|
||||
- extensions/qa-lab/src/suite.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify OpenClaw can expose plugin tools over MCP and a real MCP client can call one successfully.
|
||||
config:
|
||||
memoryFact: "MCP fact: the codename is ORBIT-9."
|
||||
query: "ORBIT-9 codename"
|
||||
expectedNeedle: "ORBIT-9"
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: serves and calls memory_search over MCP
|
||||
actions:
|
||||
- call: fs.writeFile
|
||||
args:
|
||||
- expr: "path.join(env.gateway.workspaceDir, 'MEMORY.md')"
|
||||
- expr: "`${config.memoryFact}\\n`"
|
||||
- utf8
|
||||
- call: forceMemoryIndex
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
query:
|
||||
expr: config.query
|
||||
expectedNeedle:
|
||||
expr: config.expectedNeedle
|
||||
- call: callPluginToolsMcp
|
||||
saveAs: result
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
toolName: memory_search
|
||||
args:
|
||||
query:
|
||||
expr: config.query
|
||||
maxResults: 3
|
||||
- set: text
|
||||
value:
|
||||
expr: "JSON.stringify(result.content ?? [])"
|
||||
- assert:
|
||||
expr: "text.includes(config.expectedNeedle)"
|
||||
message:
|
||||
expr: "`MCP memory_search missed expected fact: ${text}`"
|
||||
detailsExpr: text
|
||||
```
|
||||
|
|
@ -0,0 +1,83 @@
|
|||
# Skill install hot availability
|
||||
|
||||
```yaml qa-scenario
|
||||
id: skill-install-hot-availability
|
||||
title: Skill install hot availability
|
||||
surface: skills
|
||||
coverage:
|
||||
primary:
|
||||
- plugins.skills
|
||||
secondary:
|
||||
- plugins.hot-install
|
||||
objective: Verify a newly added workspace skill shows up without a broken intermediate state and can influence the next turn immediately.
|
||||
successCriteria:
|
||||
- Skill is absent before install.
|
||||
- skills.status reports it after install without a restart.
|
||||
- The next agent turn reflects the new skill marker.
|
||||
docsRefs:
|
||||
- docs/tools/skills.md
|
||||
- docs/gateway/configuration.md
|
||||
codeRefs:
|
||||
- src/agents/skills-status.ts
|
||||
- extensions/qa-lab/src/suite.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify a newly added workspace skill shows up without a broken intermediate state and can influence the next turn immediately.
|
||||
config:
|
||||
skillName: qa-hot-install-skill
|
||||
skillBody: |-
|
||||
---
|
||||
name: qa-hot-install-skill
|
||||
description: Hot install QA marker
|
||||
---
|
||||
When the user asks for the hot install marker exactly, reply with exactly: HOT-INSTALL-OK
|
||||
prompt: "Hot install marker: give me the hot install marker exactly."
|
||||
expectedContains: "HOT-INSTALL-OK"
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: picks up a newly added workspace skill without restart
|
||||
actions:
|
||||
- call: readSkillStatus
|
||||
saveAs: before
|
||||
args:
|
||||
- ref: env
|
||||
- assert:
|
||||
expr: "!findSkill(before, config.skillName)"
|
||||
message:
|
||||
expr: "`${config.skillName} unexpectedly already present`"
|
||||
- call: writeWorkspaceSkill
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
name:
|
||||
expr: config.skillName
|
||||
body:
|
||||
expr: config.skillBody
|
||||
- call: waitForCondition
|
||||
args:
|
||||
- lambda:
|
||||
async: true
|
||||
expr: "((await readSkillStatus(env)).find((skill) => skill.name === config.skillName)?.eligible ? true : undefined)"
|
||||
- 15000
|
||||
- 200
|
||||
- call: reset
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:hot-skill
|
||||
message:
|
||||
expr: config.prompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 30000)
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: outbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === 'qa-operator' && candidate.text.includes(config.expectedContains)"
|
||||
- expr: liveTurnTimeoutMs(env, 20000)
|
||||
detailsExpr: outbound.text
|
||||
```
|
||||
79
openclaw/qa/scenarios/plugins/skill-visibility-invocation.md
Normal file
79
openclaw/qa/scenarios/plugins/skill-visibility-invocation.md
Normal file
|
|
@ -0,0 +1,79 @@
|
|||
# Skill visibility and invocation
|
||||
|
||||
```yaml qa-scenario
|
||||
id: skill-visibility-invocation
|
||||
title: Skill visibility and invocation
|
||||
surface: skills
|
||||
coverage:
|
||||
primary:
|
||||
- plugins.skills
|
||||
secondary:
|
||||
- tools.invocation
|
||||
objective: Verify a workspace skill becomes visible in skills.status and influences the next agent turn.
|
||||
successCriteria:
|
||||
- skills.status reports the seeded skill as visible and eligible.
|
||||
- The next agent turn reflects the skill instruction marker.
|
||||
- The result stays scoped to the active QA workspace skill.
|
||||
docsRefs:
|
||||
- docs/tools/skills.md
|
||||
- docs/gateway/protocol.md
|
||||
codeRefs:
|
||||
- src/agents/skills-status.ts
|
||||
- extensions/qa-lab/src/suite.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify a workspace skill becomes visible in skills.status and influences the next agent turn.
|
||||
config:
|
||||
skillName: qa-visible-skill
|
||||
skillBody: |-
|
||||
---
|
||||
name: qa-visible-skill
|
||||
description: Visible QA skill marker
|
||||
---
|
||||
When the user asks for the visible skill marker exactly, or explicitly asks you to use qa-visible-skill, reply with exactly: VISIBLE-SKILL-OK
|
||||
prompt: "Use qa-visible-skill now. Reply exactly with the visible skill marker and nothing else."
|
||||
expectedContains: "VISIBLE-SKILL-OK"
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: reports visible skill and applies its marker on the next turn
|
||||
actions:
|
||||
- call: writeWorkspaceSkill
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
name:
|
||||
expr: config.skillName
|
||||
body:
|
||||
expr: config.skillBody
|
||||
- call: readSkillStatus
|
||||
saveAs: skills
|
||||
args:
|
||||
- ref: env
|
||||
- set: visible
|
||||
value:
|
||||
expr: findSkill(skills, config.skillName)
|
||||
- assert:
|
||||
expr: "visible?.eligible === true && !visible?.disabled && !visible?.blockedByAllowlist"
|
||||
message:
|
||||
expr: "`skill not visible/eligible: ${JSON.stringify(visible)}`"
|
||||
- call: reset
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:visible-skill
|
||||
message:
|
||||
expr: config.prompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 30000)
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: outbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === 'qa-operator' && candidate.text.includes(config.expectedContains)"
|
||||
- expr: liveTurnTimeoutMs(env, 20000)
|
||||
detailsExpr: outbound.text
|
||||
```
|
||||
|
|
@ -0,0 +1,91 @@
|
|||
# Approval turn tool followthrough
|
||||
|
||||
```yaml qa-scenario
|
||||
id: approval-turn-tool-followthrough
|
||||
title: Approval turn tool followthrough
|
||||
surface: harness
|
||||
coverage:
|
||||
primary:
|
||||
- runtime.approvals
|
||||
secondary:
|
||||
- tools.followthrough
|
||||
objective: Verify a short approval like "ok do it" triggers immediate tool use instead of fake-progress narration.
|
||||
successCriteria:
|
||||
- Agent can keep the pre-action turn brief.
|
||||
- The short approval leads to a real tool call on the next turn.
|
||||
- Final answer uses tool-derived evidence instead of placeholder progress text.
|
||||
docsRefs:
|
||||
- docs/help/testing.md
|
||||
- docs/channels/qa-channel.md
|
||||
codeRefs:
|
||||
- extensions/qa-lab/src/suite.ts
|
||||
- extensions/qa-lab/src/mock-openai-server.ts
|
||||
- src/agents/pi-embedded-runner/run/incomplete-turn.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify a short approval like "ok do it" triggers immediate tool use instead of fake-progress narration.
|
||||
config:
|
||||
preActionPrompt: Before acting, tell me the single file you would start with in six words or fewer. Do not use tools yet.
|
||||
approvalPrompt: ok do it. read `QA_KICKOFF_TASK.md` now and reply with the QA mission in one short sentence.
|
||||
expectedReplyAny:
|
||||
- qa
|
||||
- mission
|
||||
- testing
|
||||
- repo
|
||||
- worked
|
||||
- failed
|
||||
- blocked
|
||||
- chat flows
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: turns short approval into a real file read
|
||||
actions:
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- call: reset
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:approval-followthrough
|
||||
message:
|
||||
expr: config.preActionPrompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 20000)
|
||||
- call: waitForOutboundMessage
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === 'qa-operator'"
|
||||
- expr: liveTurnTimeoutMs(env, 20000)
|
||||
- set: beforeApprovalCursor
|
||||
value:
|
||||
expr: state.getSnapshot().messages.length
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:approval-followthrough
|
||||
message:
|
||||
expr: config.approvalPrompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 30000)
|
||||
- set: expectedReplyAny
|
||||
value:
|
||||
expr: config.expectedReplyAny.map(normalizeLowercaseStringOrEmpty)
|
||||
- call: waitForCondition
|
||||
saveAs: outbound
|
||||
args:
|
||||
- lambda:
|
||||
expr: "state.getSnapshot().messages.slice(beforeApprovalCursor).filter((candidate) => candidate.direction === 'outbound' && candidate.conversation.id === 'qa-operator' && expectedReplyAny.some((needle) => normalizeLowercaseStringOrEmpty(candidate.text).includes(needle))).at(-1)"
|
||||
- expr: liveTurnTimeoutMs(env, 20000)
|
||||
- expr: "env.providerMode === 'mock-openai' ? 100 : 250"
|
||||
- assert:
|
||||
expr: "!env.mock || ([...(await fetchJson(`${env.mock.baseUrl}/debug/requests`))].toReversed().find((request) => String(request.allInputText ?? '').includes('ok do it.') && !request.toolOutput)?.plannedToolName === 'read')"
|
||||
message:
|
||||
expr: "`expected read after approval, got ${String(([...(await fetchJson(`${env.mock.baseUrl}/debug/requests`))].toReversed().find((request) => String(request.allInputText ?? '').includes('ok do it.') && !request.toolOutput)?.plannedToolName ?? ''))}`"
|
||||
detailsExpr: outbound.text
|
||||
```
|
||||
103
openclaw/qa/scenarios/runtime/compaction-retry-mutating-tool.md
Normal file
103
openclaw/qa/scenarios/runtime/compaction-retry-mutating-tool.md
Normal file
|
|
@ -0,0 +1,103 @@
|
|||
# Compaction retry after mutating tool
|
||||
|
||||
```yaml qa-scenario
|
||||
id: compaction-retry-mutating-tool
|
||||
title: Compaction retry after mutating tool
|
||||
surface: runtime
|
||||
coverage:
|
||||
primary:
|
||||
- runtime.compaction
|
||||
secondary:
|
||||
- runtime.retry-policy
|
||||
objective: Verify a real mutating tool step keeps replay-unsafety explicit instead of disappearing into a clean-looking success if the run compacts or retries.
|
||||
successCriteria:
|
||||
- Agent reads the seeded large context before it writes.
|
||||
- Agent performs a real mutating write in the workspace.
|
||||
- Final reply explicitly says replay is unsafe after the write instead of implying replay-safe completion.
|
||||
- Scenario details preserve the observed compaction count for review context.
|
||||
docsRefs:
|
||||
- docs/help/testing.md
|
||||
- docs/help/gpt54-codex-agentic-parity.md
|
||||
codeRefs:
|
||||
- extensions/qa-lab/src/suite.ts
|
||||
- extensions/qa-lab/src/mock-openai-server.ts
|
||||
- src/agents/pi-embedded-subscribe.ts
|
||||
- src/agents/pi-embedded-subscribe.handlers.lifecycle.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify a mutating tool step keeps replay-unsafety explicit through compaction or retry pressure.
|
||||
config:
|
||||
contextFile: COMPACTION_RETRY_CONTEXT.md
|
||||
outputFile: compaction-retry-summary.txt
|
||||
promptSnippet: Compaction retry mutating tool check
|
||||
replayUnsafeNeedle: replay unsafe after write
|
||||
prompt: |-
|
||||
Compaction retry mutating tool check: read `COMPACTION_RETRY_CONTEXT.md`, then create `compaction-retry-summary.txt` in the workspace containing exactly this line:
|
||||
Replay safety: unsafe after write.
|
||||
|
||||
After the write, reply with one short sentence that includes the exact phrase `replay unsafe after write`.
|
||||
If the run gets compacted or retried, keep that replay-safety truth explicit and do not claim the task stayed replay-safe.
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: keeps replay-unsafety explicit after a mutating write under compaction pressure
|
||||
actions:
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- call: reset
|
||||
- call: fs.writeFile
|
||||
args:
|
||||
- expr: "path.join(env.gateway.workspaceDir, config.contextFile)"
|
||||
- expr: "Array.from({ length: env.providerMode === 'mock-openai' ? 160 : 2200 }, (_, index) => `compaction retry evidence block ${String(index).padStart(4, '0')} :: keep replay truth explicit after side effects`).join('\\n')"
|
||||
- utf8
|
||||
- set: sessionKey
|
||||
value:
|
||||
expr: "`agent:qa:compaction-retry:${randomUUID().slice(0, 8)}`"
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey:
|
||||
ref: sessionKey
|
||||
message:
|
||||
ref: config.prompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 90000)
|
||||
- call: waitForCondition
|
||||
saveAs: outbound
|
||||
args:
|
||||
- lambda:
|
||||
expr: "state.getSnapshot().messages.filter((candidate) => candidate.direction === 'outbound' && candidate.conversation.id === 'qa-operator' && normalizeLowercaseStringOrEmpty(candidate.text).includes(config.replayUnsafeNeedle)).at(-1)"
|
||||
- expr: liveTurnTimeoutMs(env, 45000)
|
||||
- expr: "env.providerMode === 'mock-openai' ? 100 : 250"
|
||||
- call: fs.readFile
|
||||
saveAs: writtenSummary
|
||||
args:
|
||||
- expr: "path.join(env.gateway.workspaceDir, config.outputFile)"
|
||||
- utf8
|
||||
- assert:
|
||||
expr: "writtenSummary.includes('Replay safety: unsafe after write.')"
|
||||
message:
|
||||
expr: "`summary file missed replay marker: ${writtenSummary}`"
|
||||
- if:
|
||||
expr: "Boolean(env.mock)"
|
||||
then:
|
||||
- assert:
|
||||
expr: "!env.mock || ([...(await fetchJson(`${env.mock.baseUrl}/debug/requests`))].toReversed().find((request) => String(request.allInputText ?? '').includes(config.promptSnippet) && String(request.toolOutput ?? '').includes('compaction retry evidence block'))?.plannedToolName === 'write')"
|
||||
message:
|
||||
expr: "`expected write after seeded context read, got ${String(([...(await fetchJson(`${env.mock.baseUrl}/debug/requests`))].toReversed().find((request) => String(request.allInputText ?? '').includes(config.promptSnippet) && String(request.toolOutput ?? '').includes('compaction retry evidence block'))?.plannedToolName ?? '')}`"
|
||||
- call: readRawQaSessionStore
|
||||
saveAs: store
|
||||
args:
|
||||
- ref: env
|
||||
- set: sessionEntry
|
||||
value:
|
||||
expr: "store[sessionKey]"
|
||||
- assert:
|
||||
expr: "Boolean(sessionEntry)"
|
||||
message:
|
||||
expr: "`missing QA session entry for ${sessionKey}`"
|
||||
detailsExpr: "`${outbound.text}\\ncompactionCount=${String(sessionEntry?.compactionCount ?? 0)}\\nstatus=${String(sessionEntry?.status ?? 'unknown')}`"
|
||||
```
|
||||
|
|
@ -0,0 +1,86 @@
|
|||
# Empty-response recovery after replay-safe read
|
||||
|
||||
```yaml qa-scenario
|
||||
id: empty-response-recovery-replay-safe-read
|
||||
title: Empty-response recovery after replay-safe read
|
||||
surface: runtime
|
||||
coverage:
|
||||
primary:
|
||||
- runtime.empty-response-recovery
|
||||
secondary:
|
||||
- runtime.retry-policy
|
||||
objective: Verify an empty visible GPT turn after a replay-safe read auto-continues into a visible answer.
|
||||
successCriteria:
|
||||
- Scenario is mock-openai only so live lanes do not pick it up implicitly.
|
||||
- The agent performs a replay-safe read before the empty response.
|
||||
- The runtime injects the visible-answer continuation instruction after the empty turn.
|
||||
- The final visible reply contains the exact recovery marker.
|
||||
docsRefs:
|
||||
- docs/help/testing.md
|
||||
codeRefs:
|
||||
- extensions/qa-lab/src/mock-openai-server.ts
|
||||
- src/agents/pi-embedded-runner/run/incomplete-turn.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify empty OpenAI turns recover after a replay-safe read.
|
||||
config:
|
||||
requiredProvider: mock-openai
|
||||
promptSnippet: Empty response continuation QA check
|
||||
prompt: "Empty response continuation QA check: read QA_KICKOFF_TASK.md, then answer with exactly EMPTY-RECOVERED-OK."
|
||||
expectedReply: EMPTY-RECOVERED-OK
|
||||
retryNeedle: The previous attempt did not produce a user-visible answer.
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: retries an empty replay-safe read into a visible answer
|
||||
actions:
|
||||
- assert:
|
||||
expr: "env.providerMode === 'mock-openai'"
|
||||
message: this seeded scenario is mock-openai only
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- call: reset
|
||||
- set: requestCountBefore
|
||||
value:
|
||||
expr: "env.mock ? (await fetchJson(`${env.mock.baseUrl}/debug/requests`)).length : 0"
|
||||
- set: sessionKey
|
||||
value:
|
||||
expr: "`agent:qa:empty-response-recovery:${randomUUID().slice(0, 8)}`"
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey:
|
||||
ref: sessionKey
|
||||
message:
|
||||
expr: config.prompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 45000)
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: outbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === 'qa-operator' && candidate.text.includes(config.expectedReply)"
|
||||
- expr: liveTurnTimeoutMs(env, 30000)
|
||||
- assert:
|
||||
expr: "outbound.text.includes(config.expectedReply)"
|
||||
message:
|
||||
expr: "`missing empty-response recovery marker: ${outbound.text}`"
|
||||
- if:
|
||||
expr: "Boolean(env.mock)"
|
||||
then:
|
||||
- set: scenarioRequests
|
||||
value:
|
||||
expr: "(await fetchJson(`${env.mock.baseUrl}/debug/requests`)).slice(requestCountBefore)"
|
||||
- assert:
|
||||
expr: "scenarioRequests.some((request) => String(request.allInputText ?? '').includes(config.promptSnippet) && request.plannedToolName === 'read')"
|
||||
message: expected replay-safe read request in mock trace
|
||||
- assert:
|
||||
expr: "scenarioRequests.some((request) => String(request.allInputText ?? '').includes(config.retryNeedle))"
|
||||
message: expected empty-response retry instruction in mock trace
|
||||
detailsExpr: "env.mock ? `${outbound.text}\\nrequests=${String(scenarioRequests?.length ?? 0)}` : outbound.text"
|
||||
```
|
||||
|
|
@ -0,0 +1,80 @@
|
|||
# Empty-response retry budget exhausted
|
||||
|
||||
```yaml qa-scenario
|
||||
id: empty-response-retry-budget-exhausted
|
||||
title: Empty-response retry budget exhausted
|
||||
surface: runtime
|
||||
coverage:
|
||||
primary:
|
||||
- runtime.empty-response-recovery
|
||||
secondary:
|
||||
- runtime.retry-policy
|
||||
objective: Verify repeated empty GPT turns exhaust the retry budget after one continuation attempt.
|
||||
successCriteria:
|
||||
- Scenario is mock-openai only so live lanes do not pick it up implicitly.
|
||||
- The agent performs the replay-safe read that makes retrying allowed.
|
||||
- Mock trace shows the run reaches a terminal post-read turn without ever producing the requested success marker.
|
||||
docsRefs:
|
||||
- docs/help/testing.md
|
||||
codeRefs:
|
||||
- extensions/qa-lab/src/mock-openai-server.ts
|
||||
- src/agents/pi-embedded-runner/run/incomplete-turn.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify empty-response retry exhaustion still surfaces a visible failure.
|
||||
config:
|
||||
requiredProvider: mock-openai
|
||||
promptSnippet: Empty response exhaustion QA check
|
||||
prompt: "Empty response exhaustion QA check: read QA_KICKOFF_TASK.md, then answer with exactly EMPTY-EXHAUSTED-OK."
|
||||
retryNeedle: The previous attempt did not produce a user-visible answer.
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: surfaces a retry error after empty-response exhaustion
|
||||
actions:
|
||||
- assert:
|
||||
expr: "env.providerMode === 'mock-openai'"
|
||||
message: this seeded scenario is mock-openai only
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- call: reset
|
||||
- set: requestCountBefore
|
||||
value:
|
||||
expr: "env.mock ? (await fetchJson(`${env.mock.baseUrl}/debug/requests`)).length : 0"
|
||||
- set: sessionKey
|
||||
value:
|
||||
expr: "`agent:qa:empty-response-exhausted:${randomUUID().slice(0, 8)}`"
|
||||
- call: startAgentRun
|
||||
saveAs: started
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey:
|
||||
ref: sessionKey
|
||||
message:
|
||||
expr: config.prompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 45000)
|
||||
- set: waited
|
||||
value:
|
||||
expr: "await env.gateway.call('agent.wait', { runId: started.runId, timeoutMs: liveTurnTimeoutMs(env, 45000) }, { timeoutMs: liveTurnTimeoutMs(env, 50000) })"
|
||||
- assert:
|
||||
expr: "waited?.status === 'ok'"
|
||||
message:
|
||||
expr: "`agent.wait returned ${String(waited?.status ?? 'unknown')}: ${String(waited?.error ?? '')}`"
|
||||
- if:
|
||||
expr: "Boolean(env.mock)"
|
||||
then:
|
||||
- set: scenarioRequests
|
||||
value:
|
||||
expr: "(await fetchJson(`${env.mock.baseUrl}/debug/requests`)).slice(requestCountBefore)"
|
||||
- assert:
|
||||
expr: "scenarioRequests.some((request) => String(request.allInputText ?? '').includes(config.promptSnippet) && request.plannedToolName === 'read')"
|
||||
message: expected replay-safe read request in mock trace
|
||||
- assert:
|
||||
expr: "scenarioRequests.length >= 2"
|
||||
message: expected at least the replay-safe read request and one terminal post-read turn
|
||||
detailsExpr: "env.mock ? `requests=${String(scenarioRequests?.length ?? 0)}` : String(waited?.status ?? '')"
|
||||
```
|
||||
|
|
@ -0,0 +1,95 @@
|
|||
# Reasoning-only no-auto-retry after write
|
||||
|
||||
```yaml qa-scenario
|
||||
id: reasoning-only-no-auto-retry-after-write
|
||||
title: Reasoning-only no-auto-retry after write
|
||||
surface: runtime
|
||||
coverage:
|
||||
primary:
|
||||
- runtime.reasoning-only-recovery
|
||||
secondary:
|
||||
- runtime.retry-policy
|
||||
objective: Verify a GPT-style reasoning-only turn after a mutating write stays replay-unsafe and does not auto-retry.
|
||||
successCriteria:
|
||||
- Scenario is mock-openai only so live lanes do not pick it up implicitly.
|
||||
- The agent performs the seeded mutating write.
|
||||
- Mock trace does not include an automatic reasoning-only retry instruction.
|
||||
- Mock trace stops after the write-side reasoning-only terminal turn instead of attempting a continuation.
|
||||
docsRefs:
|
||||
- docs/help/testing.md
|
||||
- docs/help/gpt54-codex-agentic-parity.md
|
||||
codeRefs:
|
||||
- extensions/qa-lab/src/mock-openai-server.ts
|
||||
- src/agents/pi-embedded-runner/run/incomplete-turn.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify reasoning-only turns after a write do not auto-retry.
|
||||
config:
|
||||
requiredProvider: mock-openai
|
||||
promptSnippet: Reasoning-only after write safety check
|
||||
prompt: "Reasoning-only after write safety check: write reasoning-only-side-effect.txt, then answer with exactly SIDE-EFFECT-GUARD-OK."
|
||||
retryNeedle: recorded reasoning but did not produce a user-visible answer
|
||||
outputFile: reasoning-only-side-effect.txt
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: keeps replay-unsafety explicit after a mutating write
|
||||
actions:
|
||||
- assert:
|
||||
expr: "env.providerMode === 'mock-openai'"
|
||||
message: this seeded scenario is mock-openai only
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- call: reset
|
||||
- set: requestCountBefore
|
||||
value:
|
||||
expr: "env.mock ? (await fetchJson(`${env.mock.baseUrl}/debug/requests`)).length : 0"
|
||||
- set: sessionKey
|
||||
value:
|
||||
expr: "`agent:qa:reasoning-only-write:${randomUUID().slice(0, 8)}`"
|
||||
- call: startAgentRun
|
||||
saveAs: started
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey:
|
||||
ref: sessionKey
|
||||
message:
|
||||
expr: config.prompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 45000)
|
||||
- set: waited
|
||||
value:
|
||||
expr: "await env.gateway.call('agent.wait', { runId: started.runId, timeoutMs: liveTurnTimeoutMs(env, 45000) }, { timeoutMs: liveTurnTimeoutMs(env, 50000) })"
|
||||
- assert:
|
||||
expr: "waited?.status === 'ok'"
|
||||
message:
|
||||
expr: "`agent.wait returned ${String(waited?.status ?? 'unknown')}: ${String(waited?.error ?? '')}`"
|
||||
- call: fs.readFile
|
||||
saveAs: sideEffect
|
||||
args:
|
||||
- expr: "path.join(env.gateway.workspaceDir, config.outputFile)"
|
||||
- utf8
|
||||
- assert:
|
||||
expr: "sideEffect.includes('side effects already happened')"
|
||||
message:
|
||||
expr: "`side-effect file missing expected contents: ${sideEffect}`"
|
||||
- if:
|
||||
expr: "Boolean(env.mock)"
|
||||
then:
|
||||
- set: scenarioRequests
|
||||
value:
|
||||
expr: "(await fetchJson(`${env.mock.baseUrl}/debug/requests`)).slice(requestCountBefore)"
|
||||
- assert:
|
||||
expr: "scenarioRequests.some((request) => String(request.allInputText ?? '').includes(config.promptSnippet) && request.plannedToolName === 'write')"
|
||||
message: expected mutating write request in mock trace
|
||||
- assert:
|
||||
expr: "!scenarioRequests.some((request) => String(request.allInputText ?? '').includes(config.retryNeedle))"
|
||||
message: reasoning-only retry instruction should not be injected after a write
|
||||
- assert:
|
||||
expr: "scenarioRequests.filter((request) => String(request.allInputText ?? '').includes(config.promptSnippet)).length === 2"
|
||||
message: expected exactly the write request plus the reasoning-only terminal request
|
||||
detailsExpr: "env.mock ? `requests=${String(scenarioRequests?.length ?? 0)} sideEffect=${sideEffect.trim()}` : sideEffect"
|
||||
```
|
||||
|
|
@ -0,0 +1,86 @@
|
|||
# Reasoning-only recovery after replay-safe read
|
||||
|
||||
```yaml qa-scenario
|
||||
id: reasoning-only-recovery-replay-safe-read
|
||||
title: Reasoning-only recovery after replay-safe read
|
||||
surface: runtime
|
||||
coverage:
|
||||
primary:
|
||||
- runtime.reasoning-only-recovery
|
||||
secondary:
|
||||
- runtime.retry-policy
|
||||
objective: Verify a GPT-style reasoning-only turn after a replay-safe read auto-continues into a visible answer.
|
||||
successCriteria:
|
||||
- Scenario is mock-openai only so live lanes do not pick it up implicitly.
|
||||
- The agent performs a replay-safe read before the reasoning-only turn.
|
||||
- The runtime injects the visible-answer continuation instruction after the reasoning-only turn.
|
||||
- The final visible reply contains the exact recovery marker.
|
||||
docsRefs:
|
||||
- docs/help/testing.md
|
||||
codeRefs:
|
||||
- extensions/qa-lab/src/mock-openai-server.ts
|
||||
- src/agents/pi-embedded-runner/run/incomplete-turn.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify reasoning-only OpenAI turns recover after a replay-safe read.
|
||||
config:
|
||||
requiredProvider: mock-openai
|
||||
promptSnippet: Reasoning-only continuation QA check
|
||||
prompt: "Reasoning-only continuation QA check: read QA_KICKOFF_TASK.md, then answer with exactly REASONING-RECOVERED-OK."
|
||||
expectedReply: REASONING-RECOVERED-OK
|
||||
retryNeedle: recorded reasoning but did not produce a user-visible answer
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: retries a replay-safe read into a visible answer
|
||||
actions:
|
||||
- assert:
|
||||
expr: "env.providerMode === 'mock-openai'"
|
||||
message: this seeded scenario is mock-openai only
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- call: reset
|
||||
- set: requestCountBefore
|
||||
value:
|
||||
expr: "env.mock ? (await fetchJson(`${env.mock.baseUrl}/debug/requests`)).length : 0"
|
||||
- set: sessionKey
|
||||
value:
|
||||
expr: "`agent:qa:reasoning-only-recovery:${randomUUID().slice(0, 8)}`"
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey:
|
||||
ref: sessionKey
|
||||
message:
|
||||
expr: config.prompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 45000)
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: outbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === 'qa-operator' && candidate.text.includes(config.expectedReply)"
|
||||
- expr: liveTurnTimeoutMs(env, 30000)
|
||||
- assert:
|
||||
expr: "outbound.text.includes(config.expectedReply)"
|
||||
message:
|
||||
expr: "`missing recovery marker: ${outbound.text}`"
|
||||
- if:
|
||||
expr: "Boolean(env.mock)"
|
||||
then:
|
||||
- set: scenarioRequests
|
||||
value:
|
||||
expr: "(await fetchJson(`${env.mock.baseUrl}/debug/requests`)).slice(requestCountBefore)"
|
||||
- assert:
|
||||
expr: "scenarioRequests.some((request) => String(request.allInputText ?? '').includes(config.promptSnippet) && request.plannedToolName === 'read')"
|
||||
message: expected replay-safe read request in mock trace
|
||||
- assert:
|
||||
expr: "scenarioRequests.some((request) => String(request.allInputText ?? '').includes(config.retryNeedle))"
|
||||
message: expected reasoning-only retry instruction in mock trace
|
||||
detailsExpr: "env.mock ? `${outbound.text}\\nrequests=${String(scenarioRequests?.length ?? 0)}` : outbound.text"
|
||||
```
|
||||
106
openclaw/qa/scenarios/runtime/runtime-inventory-drift-check.md
Normal file
106
openclaw/qa/scenarios/runtime/runtime-inventory-drift-check.md
Normal file
|
|
@ -0,0 +1,106 @@
|
|||
# Runtime inventory drift check
|
||||
|
||||
```yaml qa-scenario
|
||||
id: runtime-inventory-drift-check
|
||||
title: Runtime inventory drift check
|
||||
surface: inventory
|
||||
coverage:
|
||||
primary:
|
||||
- runtime.inventory
|
||||
objective: Verify tools.effective and skills.status stay aligned with runtime behavior after config changes.
|
||||
successCriteria:
|
||||
- Enabled tool appears before the config change.
|
||||
- After config change, disabled tool disappears from tools.effective.
|
||||
- Disabled skill appears in skills.status with disabled state.
|
||||
docsRefs:
|
||||
- docs/gateway/protocol.md
|
||||
- docs/tools/skills.md
|
||||
- docs/tools/index.md
|
||||
codeRefs:
|
||||
- src/gateway/server-methods/tools-effective.ts
|
||||
- src/gateway/server-methods/skills.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify tools.effective and skills.status stay aligned with runtime behavior after config changes.
|
||||
config:
|
||||
skillName: qa-drift-skill
|
||||
successMarker: DRIFT-SKILL-OK
|
||||
skillBody: |-
|
||||
---
|
||||
name: qa-drift-skill
|
||||
description: Drift skill marker
|
||||
---
|
||||
When the user asks for the drift skill marker exactly, reply with exactly: DRIFT-SKILL-OK
|
||||
deniedTool: image_generate
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: keeps tools.effective and skills.status aligned after config changes
|
||||
actions:
|
||||
- call: writeWorkspaceSkill
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
name:
|
||||
expr: config.skillName
|
||||
body:
|
||||
expr: config.skillBody
|
||||
- call: createSession
|
||||
saveAs: sessionKey
|
||||
args:
|
||||
- ref: env
|
||||
- Inventory drift
|
||||
- call: readEffectiveTools
|
||||
saveAs: beforeTools
|
||||
args:
|
||||
- ref: env
|
||||
- ref: sessionKey
|
||||
- assert:
|
||||
expr: "beforeTools.has(config.deniedTool)"
|
||||
message:
|
||||
expr: "`expected ${config.deniedTool} before drift patch`"
|
||||
- call: readSkillStatus
|
||||
saveAs: beforeSkills
|
||||
args:
|
||||
- ref: env
|
||||
- assert:
|
||||
expr: "Boolean(findSkill(beforeSkills, config.skillName)?.eligible)"
|
||||
message:
|
||||
expr: "`expected ${config.skillName} to be eligible before patch`"
|
||||
- call: patchConfig
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
patch:
|
||||
tools:
|
||||
deny:
|
||||
- expr: config.deniedTool
|
||||
skills:
|
||||
entries:
|
||||
expr: "({ [config.skillName]: { enabled: false } })"
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- call: readEffectiveTools
|
||||
saveAs: afterTools
|
||||
args:
|
||||
- ref: env
|
||||
- ref: sessionKey
|
||||
- assert:
|
||||
expr: "!afterTools.has(config.deniedTool)"
|
||||
message:
|
||||
expr: "`${config.deniedTool} still present after deny patch`"
|
||||
- call: readSkillStatus
|
||||
saveAs: afterSkills
|
||||
args:
|
||||
- ref: env
|
||||
- set: driftSkill
|
||||
value:
|
||||
expr: "findSkill(afterSkills, config.skillName)"
|
||||
- assert:
|
||||
expr: "Boolean(driftSkill?.disabled)"
|
||||
message:
|
||||
expr: "`expected disabled drift skill, got ${JSON.stringify(driftSkill)}`"
|
||||
detailsExpr: "`${config.deniedTool} removed, ${config.skillName} marker=${config.successMarker} disabled=${String(driftSkill.disabled)}`"
|
||||
```
|
||||
|
|
@ -0,0 +1,158 @@
|
|||
# Cron natural fire no duplicate
|
||||
|
||||
```yaml qa-scenario
|
||||
id: cron-natural-fire-no-duplicate
|
||||
title: Cron natural fire no duplicate
|
||||
surface: cron
|
||||
coverage:
|
||||
primary:
|
||||
- scheduling.cron
|
||||
secondary:
|
||||
- channels.qa-channel
|
||||
- scheduling.dedup
|
||||
objective: Verify one naturally fired cron run in a single gateway uptime produces exactly one qa-channel delivery for its marker.
|
||||
successCriteria:
|
||||
- A one-shot cron job fires from the scheduler timer without cron.run force mode.
|
||||
- The qa-channel receives exactly one outbound reply containing the run marker.
|
||||
- No second outbound reply with the same marker appears during the duplicate window.
|
||||
docsRefs:
|
||||
- docs/help/testing.md
|
||||
- docs/channels/qa-channel.md
|
||||
codeRefs:
|
||||
- src/cron/service.ts
|
||||
- src/cron/service/timer.ts
|
||||
- src/cron/run-log.ts
|
||||
- extensions/qa-lab/src/cron-run-wait.ts
|
||||
- extensions/qa-lab/src/suite-runtime-transport.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Let one cron job fire from the natural scheduler timer and assert qa-channel does not receive a duplicate delivery for the same marker.
|
||||
config:
|
||||
channelId: qa-room
|
||||
channelTitle: QA Room
|
||||
fireDelayMs: 12000
|
||||
duplicateWindowMs: 8000
|
||||
reminderPromptTemplate: "A natural QA cron dedupe check fired. Send a one-line ping back to the room containing this exact marker: {{marker}}"
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: creates a near-future cron job and waits for the scheduler timer
|
||||
actions:
|
||||
- call: reset
|
||||
- set: runStartedAt
|
||||
value:
|
||||
expr: "Date.now()"
|
||||
- set: scheduledFor
|
||||
value:
|
||||
expr: "new Date(runStartedAt + config.fireDelayMs).toISOString()"
|
||||
- set: cronMarker
|
||||
value:
|
||||
expr: "`QA-CRON-NATURAL-DEDUPE-${randomUUID().slice(0, 8)}`"
|
||||
- call: env.gateway.call
|
||||
saveAs: response
|
||||
args:
|
||||
- cron.add
|
||||
- name:
|
||||
expr: "`qa-natural-dedupe-${randomUUID()}`"
|
||||
enabled: true
|
||||
schedule:
|
||||
kind: at
|
||||
at:
|
||||
ref: scheduledFor
|
||||
sessionTarget: isolated
|
||||
wakeMode: now
|
||||
payload:
|
||||
kind: agentTurn
|
||||
message:
|
||||
expr: "config.reminderPromptTemplate.replace('{{marker}}', cronMarker)"
|
||||
delivery:
|
||||
mode: announce
|
||||
channel: qa-channel
|
||||
to:
|
||||
expr: "`channel:${config.channelId}`"
|
||||
- timeoutMs: 30000
|
||||
- set: jobId
|
||||
value:
|
||||
expr: response.id
|
||||
- assert:
|
||||
expr: "Boolean(jobId)"
|
||||
message: missing cron job id
|
||||
- set: scheduledAtMs
|
||||
value:
|
||||
expr: "new Date(response.schedule?.at ?? scheduledFor).getTime()"
|
||||
- set: scheduleDeltaMs
|
||||
value:
|
||||
expr: "scheduledAtMs - runStartedAt"
|
||||
- assert:
|
||||
expr: "scheduleDeltaMs >= config.fireDelayMs - 2000 && scheduleDeltaMs <= config.fireDelayMs + 5000"
|
||||
message:
|
||||
expr: "`expected near-future natural fire, got ${scheduleDeltaMs}ms`"
|
||||
- call: waitForCronRunCompletion
|
||||
saveAs: completedRun
|
||||
args:
|
||||
- callGateway:
|
||||
expr: "env.gateway.call.bind(env.gateway)"
|
||||
jobId:
|
||||
ref: jobId
|
||||
afterTs:
|
||||
ref: runStartedAt
|
||||
timeoutMs:
|
||||
expr: "liveTurnTimeoutMs(env, Math.max(60000, config.fireDelayMs + 45000))"
|
||||
- assert:
|
||||
expr: "Date.now() >= scheduledAtMs"
|
||||
message:
|
||||
expr: "`cron completed before scheduled time ${scheduledFor}`"
|
||||
- assert:
|
||||
expr: "completedRun?.status === 'ok'"
|
||||
message:
|
||||
expr: "`expected natural cron run ok, got ${JSON.stringify(completedRun)}`"
|
||||
detailsExpr: "`job=${jobId} marker=${cronMarker} scheduled=${scheduledFor}`"
|
||||
|
||||
- name: observes exactly one qa-channel delivery for the natural run
|
||||
actions:
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: firstOutbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === config.channelId && candidate.text.includes(cronMarker)"
|
||||
- expr: liveTurnTimeoutMs(env, 45000)
|
||||
- set: firstOutboundId
|
||||
value:
|
||||
expr: firstOutbound.id
|
||||
- set: firstOutboundIndex
|
||||
value:
|
||||
expr: "getTransportSnapshot().messages.findIndex((message) => message.id === firstOutboundId)"
|
||||
- assert:
|
||||
expr: "firstOutboundIndex >= 0"
|
||||
message: first outbound message missing from qa-channel snapshot
|
||||
- call: sleep
|
||||
args:
|
||||
- expr: config.duplicateWindowMs
|
||||
- set: duplicateMatches
|
||||
value:
|
||||
expr: "getTransportSnapshot().messages.filter((message) => message.direction === 'outbound' && message.conversation.id === config.channelId && message.text.includes(cronMarker))"
|
||||
- assert:
|
||||
expr: "duplicateMatches.length === 1"
|
||||
message:
|
||||
expr: "`expected one natural outbound delivery for ${cronMarker}, saw ${duplicateMatches.length}: ${duplicateMatches.map((message) => message.text).join(' | ')}`"
|
||||
- call: env.gateway.call
|
||||
saveAs: runsPage
|
||||
args:
|
||||
- cron.runs
|
||||
- id:
|
||||
ref: jobId
|
||||
limit: 10
|
||||
sortDir: desc
|
||||
- timeoutMs: 30000
|
||||
- set: completedRuns
|
||||
value:
|
||||
expr: "runsPage.entries.filter((entry) => entry.ts >= runStartedAt && ['ok', 'error', 'skipped'].includes(entry.status))"
|
||||
- assert:
|
||||
expr: "completedRuns.length === 1"
|
||||
message:
|
||||
expr: "`expected one completed natural cron run for ${jobId}, saw ${completedRuns.length}: ${JSON.stringify(completedRuns)}`"
|
||||
detailsExpr: "`first outbound=${firstOutboundId}; duplicate window=${config.duplicateWindowMs}ms`"
|
||||
```
|
||||
117
openclaw/qa/scenarios/scheduling/cron-one-minute-ping.md
Normal file
117
openclaw/qa/scenarios/scheduling/cron-one-minute-ping.md
Normal file
|
|
@ -0,0 +1,117 @@
|
|||
# Cron one-minute ping
|
||||
|
||||
```yaml qa-scenario
|
||||
id: cron-one-minute-ping
|
||||
title: Cron one-minute ping
|
||||
surface: cron
|
||||
coverage:
|
||||
primary:
|
||||
- scheduling.cron
|
||||
secondary:
|
||||
- channels.qa-channel
|
||||
objective: Verify the agent can schedule a cron reminder one minute in the future and receive the follow-up in the QA channel.
|
||||
successCriteria:
|
||||
- Agent schedules a cron reminder roughly one minute ahead.
|
||||
- Reminder returns through qa-channel.
|
||||
- Agent recognizes the reminder as part of the original task.
|
||||
docsRefs:
|
||||
- docs/help/testing.md
|
||||
- docs/channels/qa-channel.md
|
||||
codeRefs:
|
||||
- extensions/qa-lab/src/bus-server.ts
|
||||
- extensions/qa-lab/src/self-check.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify the agent can schedule a cron reminder one minute in the future and receive the follow-up in the QA channel.
|
||||
config:
|
||||
channelId: qa-room
|
||||
channelTitle: QA Room
|
||||
reminderPromptTemplate: "A QA cron just fired. Send a one-line ping back to the room containing this exact marker: {{marker}}"
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: stores a reminder roughly one minute ahead
|
||||
actions:
|
||||
- call: reset
|
||||
- set: at
|
||||
value:
|
||||
expr: "new Date(Date.now() + 60000).toISOString()"
|
||||
- set: cronMarker
|
||||
value:
|
||||
expr: "`QA-CRON-${randomUUID().slice(0, 8)}`"
|
||||
- call: env.gateway.call
|
||||
saveAs: response
|
||||
args:
|
||||
- cron.add
|
||||
- name:
|
||||
expr: "`qa-suite-${randomUUID()}`"
|
||||
enabled: true
|
||||
schedule:
|
||||
kind: at
|
||||
at:
|
||||
ref: at
|
||||
sessionTarget: isolated
|
||||
wakeMode: now
|
||||
payload:
|
||||
kind: agentTurn
|
||||
message:
|
||||
expr: "config.reminderPromptTemplate.replace('{{marker}}', cronMarker)"
|
||||
delivery:
|
||||
mode: announce
|
||||
channel: qa-channel
|
||||
to:
|
||||
expr: "`channel:${config.channelId}`"
|
||||
- set: scheduledAt
|
||||
value:
|
||||
expr: "response.schedule?.at ?? at"
|
||||
- set: delta
|
||||
value:
|
||||
expr: "new Date(scheduledAt).getTime() - Date.now()"
|
||||
- assert:
|
||||
expr: "delta >= 45000 && delta <= 75000"
|
||||
message:
|
||||
expr: "`expected ~1 minute schedule, got ${delta}ms`"
|
||||
- set: jobId
|
||||
value:
|
||||
expr: response.id
|
||||
detailsExpr: scheduledAt
|
||||
|
||||
- name: forces the reminder through QA channel delivery
|
||||
actions:
|
||||
- assert:
|
||||
expr: "Boolean(jobId)"
|
||||
message: missing cron job id
|
||||
- assert:
|
||||
expr: "Boolean(cronMarker)"
|
||||
message: missing cron marker
|
||||
- set: runStartedAt
|
||||
value:
|
||||
expr: "Date.now()"
|
||||
- call: env.gateway.call
|
||||
args:
|
||||
- cron.run
|
||||
- id:
|
||||
ref: jobId
|
||||
mode: force
|
||||
- timeoutMs: 30000
|
||||
- call: waitForCronRunCompletion
|
||||
args:
|
||||
- callGateway:
|
||||
expr: "env.gateway.call.bind(env.gateway)"
|
||||
jobId:
|
||||
ref: jobId
|
||||
afterTs:
|
||||
ref: runStartedAt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 45000)
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: outbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === config.channelId && candidate.text.includes(cronMarker)"
|
||||
- expr: liveTurnTimeoutMs(env, 45000)
|
||||
detailsExpr: outbound.text
|
||||
```
|
||||
153
openclaw/qa/scenarios/scheduling/cron-single-run-no-duplicate.md
Normal file
153
openclaw/qa/scenarios/scheduling/cron-single-run-no-duplicate.md
Normal file
|
|
@ -0,0 +1,153 @@
|
|||
# Cron single run no duplicate
|
||||
|
||||
```yaml qa-scenario
|
||||
id: cron-single-run-no-duplicate
|
||||
title: Cron single run no duplicate
|
||||
surface: cron
|
||||
coverage:
|
||||
primary:
|
||||
- scheduling.cron
|
||||
secondary:
|
||||
- channels.qa-channel
|
||||
- scheduling.dedup
|
||||
objective: Verify one forced cron run produces exactly one qa-channel delivery for its marker.
|
||||
successCriteria:
|
||||
- A single forced cron run completes successfully.
|
||||
- The qa-channel receives exactly one outbound reply containing the run marker.
|
||||
- No second outbound reply with the same marker appears during the duplicate window.
|
||||
docsRefs:
|
||||
- docs/help/testing.md
|
||||
- docs/channels/qa-channel.md
|
||||
codeRefs:
|
||||
- src/cron/service.ts
|
||||
- src/cron/run-log.ts
|
||||
- extensions/qa-lab/src/cron-run-wait.ts
|
||||
- extensions/qa-lab/src/suite-runtime-transport.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Force one cron run and assert qa-channel does not receive a duplicate delivery for the same marker.
|
||||
config:
|
||||
channelId: qa-room
|
||||
channelTitle: QA Room
|
||||
duplicateWindowMs: 8000
|
||||
reminderPromptTemplate: "A QA cron dedupe check fired. Send a one-line ping back to the room containing this exact marker: {{marker}}"
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: creates a future cron job and forces one run
|
||||
actions:
|
||||
- call: reset
|
||||
- set: scheduledFor
|
||||
value:
|
||||
expr: "new Date(Date.now() + 10 * 60 * 1000).toISOString()"
|
||||
- set: cronMarker
|
||||
value:
|
||||
expr: "`QA-CRON-DEDUPE-${randomUUID().slice(0, 8)}`"
|
||||
- call: env.gateway.call
|
||||
saveAs: response
|
||||
args:
|
||||
- cron.add
|
||||
- name:
|
||||
expr: "`qa-dedupe-${randomUUID()}`"
|
||||
enabled: true
|
||||
schedule:
|
||||
kind: at
|
||||
at:
|
||||
ref: scheduledFor
|
||||
sessionTarget: isolated
|
||||
wakeMode: now
|
||||
payload:
|
||||
kind: agentTurn
|
||||
message:
|
||||
expr: "config.reminderPromptTemplate.replace('{{marker}}', cronMarker)"
|
||||
delivery:
|
||||
mode: announce
|
||||
channel: qa-channel
|
||||
to:
|
||||
expr: "`channel:${config.channelId}`"
|
||||
- set: jobId
|
||||
value:
|
||||
expr: response.id
|
||||
- assert:
|
||||
expr: "Boolean(jobId)"
|
||||
message: missing cron job id
|
||||
- set: runStartedAt
|
||||
value:
|
||||
expr: "Date.now()"
|
||||
- call: env.gateway.call
|
||||
saveAs: runResponse
|
||||
args:
|
||||
- cron.run
|
||||
- id:
|
||||
ref: jobId
|
||||
mode: force
|
||||
- timeoutMs: 30000
|
||||
- assert:
|
||||
expr: "runResponse?.ok === true && runResponse?.ran !== false"
|
||||
message:
|
||||
expr: "`expected cron.run to enqueue one run, got ${JSON.stringify(runResponse)}`"
|
||||
detailsExpr: "`job=${jobId} marker=${cronMarker}`"
|
||||
|
||||
- name: observes exactly one qa-channel delivery for that run
|
||||
actions:
|
||||
- call: waitForCronRunCompletion
|
||||
saveAs: completedRun
|
||||
args:
|
||||
- callGateway:
|
||||
expr: "env.gateway.call.bind(env.gateway)"
|
||||
jobId:
|
||||
ref: jobId
|
||||
afterTs:
|
||||
ref: runStartedAt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 45000)
|
||||
- assert:
|
||||
expr: "completedRun?.status === 'ok'"
|
||||
message:
|
||||
expr: "`expected cron run ok, got ${JSON.stringify(completedRun)}`"
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: firstOutbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === config.channelId && candidate.text.includes(cronMarker)"
|
||||
- expr: liveTurnTimeoutMs(env, 45000)
|
||||
- set: firstOutboundId
|
||||
value:
|
||||
expr: firstOutbound.id
|
||||
- set: firstOutboundIndex
|
||||
value:
|
||||
expr: "getTransportSnapshot().messages.findIndex((message) => message.id === firstOutboundId)"
|
||||
- assert:
|
||||
expr: "firstOutboundIndex >= 0"
|
||||
message: first outbound message missing from qa-channel snapshot
|
||||
- call: sleep
|
||||
args:
|
||||
- expr: config.duplicateWindowMs
|
||||
- set: duplicateMatches
|
||||
value:
|
||||
expr: "getTransportSnapshot().messages.filter((message) => message.direction === 'outbound' && message.conversation.id === config.channelId && message.text.includes(cronMarker))"
|
||||
- assert:
|
||||
expr: "duplicateMatches.length === 1"
|
||||
message:
|
||||
expr: "`expected one outbound delivery for ${cronMarker}, saw ${duplicateMatches.length}: ${duplicateMatches.map((message) => message.text).join(' | ')}`"
|
||||
- call: env.gateway.call
|
||||
saveAs: runsPage
|
||||
args:
|
||||
- cron.runs
|
||||
- id:
|
||||
ref: jobId
|
||||
limit: 10
|
||||
sortDir: desc
|
||||
- timeoutMs: 30000
|
||||
- set: completedRuns
|
||||
value:
|
||||
expr: "runsPage.entries.filter((entry) => entry.ts >= runStartedAt && ['ok', 'error', 'skipped'].includes(entry.status))"
|
||||
- assert:
|
||||
expr: "completedRuns.length === 1"
|
||||
message:
|
||||
expr: "`expected one completed cron run for ${jobId}, saw ${completedRuns.length}: ${JSON.stringify(completedRuns)}`"
|
||||
detailsExpr: "`first outbound=${firstOutboundId}; duplicate window=${config.duplicateWindowMs}ms`"
|
||||
```
|
||||
|
|
@ -0,0 +1,276 @@
|
|||
# Control UI plus qa-channel image roundtrip
|
||||
|
||||
```yaml qa-scenario
|
||||
id: control-ui-qa-channel-image-roundtrip
|
||||
title: Control UI plus qa-channel image roundtrip
|
||||
surface: control-ui
|
||||
coverage:
|
||||
primary:
|
||||
- ui.control
|
||||
secondary:
|
||||
- media.image-understanding
|
||||
- channels.qa-channel
|
||||
objective: Verify the embedded Control UI can observe a qa-channel-backed session while the fake channel injects text and image turns that the agent answers correctly.
|
||||
successCriteria:
|
||||
- Control UI opens directly on the target qa-channel session.
|
||||
- A text prompt delivered through qa-channel produces a correct outbound reply.
|
||||
- A later qa-channel image message produces a correct image-aware reply.
|
||||
- The Control UI transcript shows both transport-side prompts and both final answers.
|
||||
docsRefs:
|
||||
- docs/concepts/qa-e2e-automation.md
|
||||
- docs/channels/qa-channel.md
|
||||
codeRefs:
|
||||
- extensions/qa-lab/src/scenario-runtime-api.ts
|
||||
- extensions/qa-lab/src/suite.ts
|
||||
- extensions/qa-lab/src/web-runtime.ts
|
||||
- ui/src/ui/views/chat.ts
|
||||
gatewayRuntime:
|
||||
forwardHostHome: true
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Open the Control UI on a qa-channel session with the generic QA web driver, inject text and image turns through qa-channel, and verify the replies in both the transport log and the UI transcript.
|
||||
config:
|
||||
conversationId: control-ui-e2e
|
||||
textPrompt: "Control UI bridge check. Marker exact marker: `ui bridge armed`"
|
||||
uiExpectedNeedle: ui bridge armed
|
||||
imagePrompt: "Image understanding check: describe the top and bottom colors in the attached image in one short sentence."
|
||||
imagePromptNeedle: image understanding check
|
||||
requiredColorGroups:
|
||||
- [red, scarlet, crimson]
|
||||
- [blue, azure, teal, cyan, aqua]
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: opens control ui on the qa-channel-backed session
|
||||
actions:
|
||||
- call: reset
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- expr: liveTurnTimeoutMs(env, 60000)
|
||||
- call: waitForQaChannelReady
|
||||
args:
|
||||
- ref: env
|
||||
- expr: liveTurnTimeoutMs(env, 60000)
|
||||
- call: fetchJson
|
||||
saveAs: bootstrap
|
||||
args:
|
||||
- expr: "`${lab.baseUrl}/api/bootstrap`"
|
||||
- assert:
|
||||
expr: "Boolean(bootstrap.controlUiEmbeddedUrl)"
|
||||
message: qa-lab bootstrap did not expose controlUiEmbeddedUrl
|
||||
- set: uiSessionKey
|
||||
value:
|
||||
expr: "buildAgentSessionKey({ agentId: env.cfg.agents?.list?.find((agent) => agent.default)?.id ?? env.cfg.agents?.list?.[0]?.id ?? 'main', channel: 'qa-channel', accountId: 'default', peer: { kind: 'direct', id: config.conversationId }, dmScope: env.cfg.session?.dmScope, identityLinks: env.cfg.session?.identityLinks })"
|
||||
- set: controlUiChatUrl
|
||||
value:
|
||||
expr: "(() => { const url = new URL(String(bootstrap.controlUiEmbeddedUrl)); url.pathname = `${url.pathname.replace(/\\/$/, '')}/chat`; url.searchParams.set('session', uiSessionKey); return url.toString(); })()"
|
||||
- call: webOpenPage
|
||||
saveAs: uiTab
|
||||
args:
|
||||
- url:
|
||||
ref: controlUiChatUrl
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 60000)
|
||||
- set: uiPageId
|
||||
value:
|
||||
expr: "uiTab.pageId"
|
||||
- call: webWait
|
||||
args:
|
||||
- pageId:
|
||||
ref: uiPageId
|
||||
selector: textarea
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 45000)
|
||||
- call: waitForCondition
|
||||
saveAs: uiReadySnapshot
|
||||
args:
|
||||
- lambda:
|
||||
async: true
|
||||
expr: "await (async () => { const snapshot = await webSnapshot({ pageId: uiPageId, maxChars: 12000, timeoutMs: liveTurnTimeoutMs(env, 30000) }); const text = normalizeLowercaseStringOrEmpty(snapshot.text); return text.includes('ready to chat') ? snapshot : undefined; })()"
|
||||
- expr: liveTurnTimeoutMs(env, 45000)
|
||||
- 500
|
||||
- assert:
|
||||
expr: "Boolean(uiPageId)"
|
||||
message: control ui page was not available
|
||||
detailsExpr: "uiReadySnapshot.text"
|
||||
- name: text injected through qa-channel gets a correct transport reply
|
||||
actions:
|
||||
- set: firstInboundStartIndex
|
||||
value:
|
||||
expr: "state.getSnapshot().messages.filter((message) => message.direction === 'inbound').length"
|
||||
- set: firstOutboundStartIndex
|
||||
value:
|
||||
expr: "state.getSnapshot().messages.filter((message) => message.direction === 'outbound').length"
|
||||
- call: injectInboundMessage
|
||||
args:
|
||||
- accountId: default
|
||||
conversation:
|
||||
id:
|
||||
expr: config.conversationId
|
||||
kind: direct
|
||||
senderId:
|
||||
expr: config.conversationId
|
||||
senderName: Control UI QA
|
||||
text:
|
||||
expr: config.textPrompt
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: uiOutbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === config.conversationId && normalizeLowercaseStringOrEmpty(candidate.text).includes(config.uiExpectedNeedle)"
|
||||
- expr: liveTurnTimeoutMs(env, 45000)
|
||||
- sinceIndex:
|
||||
ref: firstOutboundStartIndex
|
||||
- call: readRawQaSessionStore
|
||||
saveAs: rawSessionStore
|
||||
args:
|
||||
- ref: env
|
||||
- set: rawSessionStoreKeys
|
||||
value:
|
||||
expr: "Object.keys(rawSessionStore)"
|
||||
detailsExpr: "`${uiOutbound.text}\\nSTORE:${JSON.stringify(rawSessionStoreKeys)}`"
|
||||
- name: text injected through qa-channel renders in a fresh control ui load
|
||||
actions:
|
||||
- call: webOpenPage
|
||||
saveAs: uiAckTab
|
||||
args:
|
||||
- url:
|
||||
ref: controlUiChatUrl
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 60000)
|
||||
- set: uiAckPageId
|
||||
value:
|
||||
expr: "uiAckTab.pageId"
|
||||
- call: webWait
|
||||
args:
|
||||
- pageId:
|
||||
ref: uiAckPageId
|
||||
selector: textarea
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 45000)
|
||||
- try:
|
||||
actions:
|
||||
- call: waitForCondition
|
||||
saveAs: uiAckSnapshot
|
||||
args:
|
||||
- lambda:
|
||||
async: true
|
||||
expr: "await (async () => { const snapshot = await webSnapshot({ pageId: uiAckPageId, maxChars: 12000, timeoutMs: liveTurnTimeoutMs(env, 30000) }); const text = normalizeLowercaseStringOrEmpty(snapshot.text); return text.includes(config.uiExpectedNeedle) && text.includes('control ui bridge check') ? snapshot : undefined; })()"
|
||||
- expr: liveTurnTimeoutMs(env, 45000)
|
||||
- 500
|
||||
catch:
|
||||
- call: webSnapshot
|
||||
saveAs: uiAckFailureSnapshot
|
||||
args:
|
||||
- pageId:
|
||||
ref: uiAckPageId
|
||||
maxChars: 12000
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 15000)
|
||||
- call: webEvaluate
|
||||
saveAs: uiAckFailureState
|
||||
args:
|
||||
- pageId:
|
||||
ref: uiAckPageId
|
||||
expression: "(() => { const app = document.querySelector('openclaw-app'); return app ? { sessionKey: app.sessionKey, settingsSessionKey: app.settings?.sessionKey, lastActiveSessionKey: app.settings?.lastActiveSessionKey, chatMessages: Array.isArray(app.chatMessages) ? app.chatMessages.length : null, chatLoading: app.chatLoading, lastError: app.lastError, connected: app.connected } : null; })()"
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 15000)
|
||||
- throw:
|
||||
expr: "`control ui text transcript missing after fresh load. state=${JSON.stringify(uiAckFailureState)} snapshot: ${uiAckFailureSnapshot.text}`"
|
||||
detailsExpr: "uiAckSnapshot.text"
|
||||
- name: image injected through qa-channel gets a correct transport reply
|
||||
actions:
|
||||
- set: secondOutboundStartIndex
|
||||
value:
|
||||
expr: "state.getSnapshot().messages.filter((message) => message.direction === 'outbound').length"
|
||||
- call: injectInboundMessage
|
||||
args:
|
||||
- accountId: default
|
||||
conversation:
|
||||
id:
|
||||
expr: config.conversationId
|
||||
kind: direct
|
||||
senderId:
|
||||
expr: config.conversationId
|
||||
senderName: Control UI QA
|
||||
text:
|
||||
expr: config.imagePrompt
|
||||
attachments:
|
||||
- kind: image
|
||||
mimeType: image/png
|
||||
fileName: red-top-blue-bottom.png
|
||||
altText: red on top blue on bottom
|
||||
contentBase64:
|
||||
expr: imageUnderstandingValidPngBase64
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: imageOutbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === config.conversationId && config.requiredColorGroups.every((group) => group.some((color) => normalizeLowercaseStringOrEmpty(candidate.text).includes(color)))"
|
||||
- expr: liveTurnTimeoutMs(env, 45000)
|
||||
- sinceIndex:
|
||||
ref: secondOutboundStartIndex
|
||||
- set: missingColorGroup
|
||||
value:
|
||||
expr: "config.requiredColorGroups.find((group) => !group.some((color) => normalizeLowercaseStringOrEmpty(imageOutbound.text).includes(color)))"
|
||||
- assert:
|
||||
expr: "!missingColorGroup"
|
||||
message:
|
||||
expr: "`missing expected colors in image reply: ${imageOutbound.text}`"
|
||||
detailsExpr: "imageOutbound.text"
|
||||
- name: image injected through qa-channel renders in a fresh control ui load
|
||||
actions:
|
||||
- call: webOpenPage
|
||||
saveAs: uiImageTab
|
||||
args:
|
||||
- url:
|
||||
ref: controlUiChatUrl
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 60000)
|
||||
- set: uiImagePageId
|
||||
value:
|
||||
expr: "uiImageTab.pageId"
|
||||
- call: webWait
|
||||
args:
|
||||
- pageId:
|
||||
ref: uiImagePageId
|
||||
selector: textarea
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 45000)
|
||||
- try:
|
||||
actions:
|
||||
- call: waitForCondition
|
||||
saveAs: uiImageSnapshot
|
||||
args:
|
||||
- lambda:
|
||||
async: true
|
||||
expr: "await (async () => { const snapshot = await webSnapshot({ pageId: uiImagePageId, maxChars: 12000, timeoutMs: liveTurnTimeoutMs(env, 30000) }); const text = normalizeLowercaseStringOrEmpty(snapshot.text); const hasPrompt = text.includes(config.imagePromptNeedle); const hasColors = config.requiredColorGroups.every((group) => group.some((color) => text.includes(color))); return hasPrompt && hasColors ? snapshot : undefined; })()"
|
||||
- expr: liveTurnTimeoutMs(env, 45000)
|
||||
- 500
|
||||
catch:
|
||||
- call: webSnapshot
|
||||
saveAs: uiImageFailureSnapshot
|
||||
args:
|
||||
- pageId:
|
||||
ref: uiImagePageId
|
||||
maxChars: 12000
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 15000)
|
||||
- call: webEvaluate
|
||||
saveAs: uiImageFailureState
|
||||
args:
|
||||
- pageId:
|
||||
ref: uiImagePageId
|
||||
expression: "(() => { const app = document.querySelector('openclaw-app'); return app ? { sessionKey: app.sessionKey, settingsSessionKey: app.settings?.sessionKey, lastActiveSessionKey: app.settings?.lastActiveSessionKey, chatMessages: Array.isArray(app.chatMessages) ? app.chatMessages.length : null, chatLoading: app.chatLoading, lastError: app.lastError, connected: app.connected } : null; })()"
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 15000)
|
||||
- throw:
|
||||
expr: "`control ui image transcript missing after fresh load. state=${JSON.stringify(uiImageFailureState)} snapshot: ${uiImageFailureSnapshot.text}`"
|
||||
detailsExpr: "uiImageSnapshot.text"
|
||||
```
|
||||
67
openclaw/qa/scenarios/workspace/lobster-invaders-build.md
Normal file
67
openclaw/qa/scenarios/workspace/lobster-invaders-build.md
Normal file
|
|
@ -0,0 +1,67 @@
|
|||
# Build Lobster Invaders
|
||||
|
||||
```yaml qa-scenario
|
||||
id: lobster-invaders-build
|
||||
title: Build Lobster Invaders
|
||||
surface: workspace
|
||||
coverage:
|
||||
primary:
|
||||
- workspace.artifacts
|
||||
secondary:
|
||||
- workspace.builds
|
||||
objective: Verify the agent can read the repo, create a tiny playable artifact, and report what changed.
|
||||
successCriteria:
|
||||
- Agent inspects source before coding.
|
||||
- Agent builds a tiny playable Lobster Invaders artifact.
|
||||
- Agent explains how to run or view the artifact.
|
||||
docsRefs:
|
||||
- docs/help/testing.md
|
||||
- docs/web/dashboard.md
|
||||
codeRefs:
|
||||
- extensions/qa-lab/src/report.ts
|
||||
- extensions/qa-lab/web/src/app.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify the agent can read the repo, create a tiny playable artifact, and report what changed.
|
||||
config:
|
||||
prompt: Read the QA kickoff context first, then build a tiny Lobster Invaders HTML game at ./lobster-invaders.html in this workspace and tell me where it is.
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: creates the artifact after reading context
|
||||
actions:
|
||||
- call: reset
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:lobster-invaders
|
||||
message:
|
||||
expr: config.prompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 30000)
|
||||
- call: waitForOutboundMessage
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === 'qa-operator'"
|
||||
- set: artifactPath
|
||||
value:
|
||||
expr: "path.join(env.gateway.workspaceDir, 'lobster-invaders.html')"
|
||||
- call: waitForCondition
|
||||
saveAs: artifact
|
||||
args:
|
||||
- lambda:
|
||||
async: true
|
||||
expr: "((await fs.readFile(artifactPath, 'utf8').catch(() => null))?.includes('Lobster Invaders') ? await fs.readFile(artifactPath, 'utf8').catch(() => null) : undefined)"
|
||||
- expr: liveTurnTimeoutMs(env, 20000)
|
||||
- 250
|
||||
- assert:
|
||||
expr: "artifact.includes('Lobster Invaders')"
|
||||
message: missing Lobster Invaders artifact
|
||||
- assert:
|
||||
expr: "!env.mock || (await fetchJson(`${env.mock.baseUrl}/debug/requests`)).some((request) => (request.toolOutput ?? '').includes('QA mission'))"
|
||||
message: expected pre-write read evidence
|
||||
detailsExpr: "'lobster-invaders.html'"
|
||||
```
|
||||
|
|
@ -0,0 +1,167 @@
|
|||
# Medium game plan Codex harness
|
||||
|
||||
```yaml qa-scenario
|
||||
id: medium-game-plan-codex-harness
|
||||
title: Medium game plan Codex harness
|
||||
surface: workspace
|
||||
coverage:
|
||||
primary:
|
||||
- workspace.planning
|
||||
secondary:
|
||||
- models.codex-cli
|
||||
objective: Verify the Codex app-server harness can plan and build a medium-complex self-contained browser game.
|
||||
successCriteria:
|
||||
- A live-frontier run fails fast unless the selected primary model is codex/gpt-5.4.
|
||||
- The scenario forces the Codex embedded harness and disables PI fallback.
|
||||
- The prompt explicitly asks the agent to enter plan mode before editing.
|
||||
- The agent writes a self-contained HTML game with a canvas loop, controls, scoring, waves, pause, and restart.
|
||||
docsRefs:
|
||||
- docs/plugins/sdk-agent-harness.md
|
||||
- docs/gateway/configuration-reference.md
|
||||
- docs/help/testing.md
|
||||
codeRefs:
|
||||
- extensions/codex/harness.ts
|
||||
- src/agents/harness/selection.ts
|
||||
- extensions/qa-lab/src/suite.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Run with `pnpm openclaw qa suite --provider-mode live-frontier --model codex/gpt-5.4 --alt-model codex/gpt-5.4 --scenario medium-game-plan-codex-harness`.
|
||||
config:
|
||||
requiredProvider: codex
|
||||
requiredModel: gpt-5.4
|
||||
harnessRuntime: codex
|
||||
harnessFallback: none
|
||||
artifactFile: star-garden-defenders-codex.html
|
||||
gameTitle: Star Garden Defenders
|
||||
minBytes: 5000
|
||||
buildPrompt: |-
|
||||
Enter plan mode first and write a short implementation plan before editing.
|
||||
|
||||
Then build a medium-complex, self-contained browser game at ./star-garden-defenders-codex.html.
|
||||
|
||||
Game: Star Garden Defenders.
|
||||
Requirements:
|
||||
- one HTML file only; no external assets, fonts, scripts, or network calls
|
||||
- canvas-based arcade loop with requestAnimationFrame
|
||||
- keyboard controls and mouse or pointer support
|
||||
- player movement, enemy waves, collectibles or power-ups, collision handling
|
||||
- score, lives or health, wave number, pause, restart, and game-over state
|
||||
- polished inline CSS and clear on-screen controls
|
||||
- after writing the file, reply with the filename and the main systems implemented
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: confirms GPT-5.4 Codex harness target
|
||||
actions:
|
||||
- set: selected
|
||||
value:
|
||||
expr: splitModelRef(env.primaryModel)
|
||||
- assert:
|
||||
expr: "env.providerMode !== 'live-frontier' || selected?.provider === config.requiredProvider"
|
||||
message:
|
||||
expr: "`expected live primary provider ${config.requiredProvider}, got ${env.primaryModel}`"
|
||||
- assert:
|
||||
expr: "env.providerMode !== 'live-frontier' || selected?.model === config.requiredModel"
|
||||
message:
|
||||
expr: "`expected live primary model ${config.requiredModel}, got ${env.primaryModel}`"
|
||||
- if:
|
||||
expr: "env.providerMode !== 'live-frontier'"
|
||||
then:
|
||||
- assert: "true"
|
||||
else:
|
||||
- call: patchConfig
|
||||
saveAs: patchResult
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
patch:
|
||||
agents:
|
||||
defaults:
|
||||
embeddedHarness:
|
||||
runtime:
|
||||
expr: config.harnessRuntime
|
||||
fallback:
|
||||
expr: config.harnessFallback
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- call: waitForQaChannelReady
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- call: readConfigSnapshot
|
||||
saveAs: snapshot
|
||||
args:
|
||||
- ref: env
|
||||
- assert:
|
||||
expr: "snapshot.config.agents?.defaults?.embeddedHarness?.runtime === config.harnessRuntime"
|
||||
message:
|
||||
expr: "`expected embeddedHarness.runtime=${config.harnessRuntime}, got ${JSON.stringify(snapshot.config.agents?.defaults?.embeddedHarness)}`"
|
||||
- assert:
|
||||
expr: "snapshot.config.agents?.defaults?.embeddedHarness?.fallback === config.harnessFallback"
|
||||
message:
|
||||
expr: "`expected embeddedHarness.fallback=${config.harnessFallback}, got ${JSON.stringify(snapshot.config.agents?.defaults?.embeddedHarness)}`"
|
||||
detailsExpr: "env.providerMode === 'live-frontier' ? `provider=${selected?.provider} model=${selected?.model} runtime=${snapshot.config.agents?.defaults?.embeddedHarness?.runtime} fallback=${snapshot.config.agents?.defaults?.embeddedHarness?.fallback}` : `mock mode: parsed ${scenario.id}`"
|
||||
- name: builds the medium game artifact
|
||||
actions:
|
||||
- if:
|
||||
expr: "env.providerMode !== 'live-frontier'"
|
||||
then:
|
||||
- assert: "true"
|
||||
else:
|
||||
- call: reset
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:medium-game-codex
|
||||
message:
|
||||
expr: config.buildPrompt
|
||||
provider:
|
||||
expr: selected?.provider
|
||||
model:
|
||||
expr: selected?.model
|
||||
timeoutMs:
|
||||
expr: resolveQaLiveTurnTimeoutMs(env, 420000, env.primaryModel)
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: outbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === 'qa-operator' && candidate.text.includes(config.artifactFile)"
|
||||
- expr: resolveQaLiveTurnTimeoutMs(env, 60000, env.primaryModel)
|
||||
- set: artifactPath
|
||||
value:
|
||||
expr: "path.join(env.gateway.workspaceDir, config.artifactFile)"
|
||||
- call: waitForCondition
|
||||
saveAs: artifact
|
||||
args:
|
||||
- lambda:
|
||||
async: true
|
||||
expr: "((await fs.readFile(artifactPath, 'utf8').catch(() => '')).includes(config.gameTitle) ? await fs.readFile(artifactPath, 'utf8').catch(() => '') : undefined)"
|
||||
- expr: resolveQaLiveTurnTimeoutMs(env, 60000, env.primaryModel)
|
||||
- 500
|
||||
- set: artifactLower
|
||||
value:
|
||||
expr: normalizeLowercaseStringOrEmpty(artifact)
|
||||
- assert:
|
||||
expr: "artifact.length >= config.minBytes"
|
||||
message:
|
||||
expr: "`expected medium game artifact >= ${config.minBytes} bytes, got ${artifact.length}`"
|
||||
- assert:
|
||||
expr: "artifactLower.includes('star garden defenders') && artifactLower.includes('<canvas') && artifactLower.includes('requestanimationframe')"
|
||||
message: missing title, canvas, or animation loop
|
||||
- assert:
|
||||
expr: "artifactLower.includes('keydown') || artifactLower.includes('keyup')"
|
||||
message: missing keyboard controls
|
||||
- assert:
|
||||
expr: "artifactLower.includes('score') && artifactLower.includes('wave') && artifactLower.includes('pause') && artifactLower.includes('restart')"
|
||||
message: missing score, wave, pause, or restart systems
|
||||
- assert:
|
||||
expr: "outbound.text.includes(config.artifactFile)"
|
||||
message:
|
||||
expr: "`final reply did not mention ${config.artifactFile}: ${outbound.text}`"
|
||||
detailsExpr: "env.providerMode !== 'live-frontier' ? 'mock mode: skipped live medium-game build' : `${config.artifactFile} bytes=${artifact.length}`"
|
||||
```
|
||||
163
openclaw/qa/scenarios/workspace/medium-game-plan-pi-harness.md
Normal file
163
openclaw/qa/scenarios/workspace/medium-game-plan-pi-harness.md
Normal file
|
|
@ -0,0 +1,163 @@
|
|||
# Medium game plan PI harness
|
||||
|
||||
```yaml qa-scenario
|
||||
id: medium-game-plan-pi-harness
|
||||
title: Medium game plan PI harness
|
||||
surface: workspace
|
||||
coverage:
|
||||
primary:
|
||||
- workspace.planning
|
||||
secondary:
|
||||
- agents.pi-harness
|
||||
objective: Verify GPT-5.4 can use the PI harness to plan and build a medium-complex self-contained browser game.
|
||||
successCriteria:
|
||||
- A live-frontier run fails fast unless the selected primary model is openai/gpt-5.4.
|
||||
- The scenario forces the embedded PI harness before the build turn.
|
||||
- The prompt explicitly asks the agent to enter plan mode before editing.
|
||||
- The agent writes a self-contained HTML game with a canvas loop, controls, scoring, waves, pause, and restart.
|
||||
docsRefs:
|
||||
- docs/plugins/sdk-agent-harness.md
|
||||
- docs/gateway/configuration-reference.md
|
||||
- docs/help/testing.md
|
||||
codeRefs:
|
||||
- src/agents/harness/selection.ts
|
||||
- src/agents/harness/builtin-pi.ts
|
||||
- extensions/qa-lab/src/suite.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Run with `pnpm openclaw qa suite --provider-mode live-frontier --model openai/gpt-5.4 --alt-model openai/gpt-5.4 --scenario medium-game-plan-pi-harness`.
|
||||
config:
|
||||
requiredProvider: openai
|
||||
requiredModel: gpt-5.4
|
||||
harnessRuntime: pi
|
||||
harnessFallback: pi
|
||||
artifactFile: star-garden-defenders-pi.html
|
||||
gameTitle: Star Garden Defenders
|
||||
minBytes: 5000
|
||||
buildPrompt: |-
|
||||
Enter plan mode first and write a short implementation plan before editing.
|
||||
|
||||
Then build a medium-complex, self-contained browser game at ./star-garden-defenders-pi.html.
|
||||
|
||||
Game: Star Garden Defenders.
|
||||
Requirements:
|
||||
- one HTML file only; no external assets, fonts, scripts, or network calls
|
||||
- canvas-based arcade loop with requestAnimationFrame
|
||||
- keyboard controls and mouse or pointer support
|
||||
- player movement, enemy waves, collectibles or power-ups, collision handling
|
||||
- score, lives or health, wave number, pause, restart, and game-over state
|
||||
- polished inline CSS and clear on-screen controls
|
||||
- after writing the file, reply with the filename and the main systems implemented
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: confirms GPT-5.4 PI harness target
|
||||
actions:
|
||||
- set: selected
|
||||
value:
|
||||
expr: splitModelRef(env.primaryModel)
|
||||
- assert:
|
||||
expr: "env.providerMode !== 'live-frontier' || selected?.provider === config.requiredProvider"
|
||||
message:
|
||||
expr: "`expected live primary provider ${config.requiredProvider}, got ${env.primaryModel}`"
|
||||
- assert:
|
||||
expr: "env.providerMode !== 'live-frontier' || selected?.model === config.requiredModel"
|
||||
message:
|
||||
expr: "`expected live primary model ${config.requiredModel}, got ${env.primaryModel}`"
|
||||
- if:
|
||||
expr: "env.providerMode !== 'live-frontier'"
|
||||
then:
|
||||
- assert: "true"
|
||||
else:
|
||||
- call: patchConfig
|
||||
saveAs: patchResult
|
||||
args:
|
||||
- env:
|
||||
ref: env
|
||||
patch:
|
||||
agents:
|
||||
defaults:
|
||||
embeddedHarness:
|
||||
runtime:
|
||||
expr: config.harnessRuntime
|
||||
fallback:
|
||||
expr: config.harnessFallback
|
||||
- call: waitForGatewayHealthy
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- call: waitForQaChannelReady
|
||||
args:
|
||||
- ref: env
|
||||
- 60000
|
||||
- call: readConfigSnapshot
|
||||
saveAs: snapshot
|
||||
args:
|
||||
- ref: env
|
||||
- assert:
|
||||
expr: "snapshot.config.agents?.defaults?.embeddedHarness?.runtime === config.harnessRuntime"
|
||||
message:
|
||||
expr: "`expected embeddedHarness.runtime=${config.harnessRuntime}, got ${JSON.stringify(snapshot.config.agents?.defaults?.embeddedHarness)}`"
|
||||
detailsExpr: "env.providerMode === 'live-frontier' ? `provider=${selected?.provider} model=${selected?.model} runtime=${snapshot.config.agents?.defaults?.embeddedHarness?.runtime}` : `mock mode: parsed ${scenario.id}`"
|
||||
- name: builds the medium game artifact
|
||||
actions:
|
||||
- if:
|
||||
expr: "env.providerMode !== 'live-frontier'"
|
||||
then:
|
||||
- assert: "true"
|
||||
else:
|
||||
- call: reset
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:medium-game-pi
|
||||
message:
|
||||
expr: config.buildPrompt
|
||||
provider:
|
||||
expr: selected?.provider
|
||||
model:
|
||||
expr: selected?.model
|
||||
timeoutMs:
|
||||
expr: resolveQaLiveTurnTimeoutMs(env, 420000, env.primaryModel)
|
||||
- call: waitForOutboundMessage
|
||||
saveAs: outbound
|
||||
args:
|
||||
- ref: state
|
||||
- lambda:
|
||||
params: [candidate]
|
||||
expr: "candidate.conversation.id === 'qa-operator' && candidate.text.includes(config.artifactFile)"
|
||||
- expr: resolveQaLiveTurnTimeoutMs(env, 60000, env.primaryModel)
|
||||
- set: artifactPath
|
||||
value:
|
||||
expr: "path.join(env.gateway.workspaceDir, config.artifactFile)"
|
||||
- call: waitForCondition
|
||||
saveAs: artifact
|
||||
args:
|
||||
- lambda:
|
||||
async: true
|
||||
expr: "((await fs.readFile(artifactPath, 'utf8').catch(() => '')).includes(config.gameTitle) ? await fs.readFile(artifactPath, 'utf8').catch(() => '') : undefined)"
|
||||
- expr: resolveQaLiveTurnTimeoutMs(env, 60000, env.primaryModel)
|
||||
- 500
|
||||
- set: artifactLower
|
||||
value:
|
||||
expr: normalizeLowercaseStringOrEmpty(artifact)
|
||||
- assert:
|
||||
expr: "artifact.length >= config.minBytes"
|
||||
message:
|
||||
expr: "`expected medium game artifact >= ${config.minBytes} bytes, got ${artifact.length}`"
|
||||
- assert:
|
||||
expr: "artifactLower.includes('star garden defenders') && artifactLower.includes('<canvas') && artifactLower.includes('requestanimationframe')"
|
||||
message: missing title, canvas, or animation loop
|
||||
- assert:
|
||||
expr: "artifactLower.includes('keydown') || artifactLower.includes('keyup')"
|
||||
message: missing keyboard controls
|
||||
- assert:
|
||||
expr: "artifactLower.includes('score') && artifactLower.includes('wave') && artifactLower.includes('pause') && artifactLower.includes('restart')"
|
||||
message: missing score, wave, pause, or restart systems
|
||||
- assert:
|
||||
expr: "outbound.text.includes(config.artifactFile)"
|
||||
message:
|
||||
expr: "`final reply did not mention ${config.artifactFile}: ${outbound.text}`"
|
||||
detailsExpr: "env.providerMode !== 'live-frontier' ? 'mock mode: skipped live medium-game build' : `${config.artifactFile} bytes=${artifact.length}`"
|
||||
```
|
||||
|
|
@ -0,0 +1,80 @@
|
|||
# Source and docs discovery report
|
||||
|
||||
```yaml qa-scenario
|
||||
id: source-docs-discovery-report
|
||||
title: Source and docs discovery report
|
||||
surface: discovery
|
||||
coverage:
|
||||
primary:
|
||||
- workspace.repo-discovery
|
||||
secondary:
|
||||
- docs.discovery
|
||||
objective: Verify the agent can read repo docs and source, expand the QA plan, and publish a worked or did-not-work report.
|
||||
successCriteria:
|
||||
- Agent reads docs and source before proposing more tests.
|
||||
- Agent identifies extra candidate scenarios beyond the seed list.
|
||||
- Agent ends with a worked or failed QA report.
|
||||
docsRefs:
|
||||
- docs/help/testing.md
|
||||
- docs/web/dashboard.md
|
||||
- docs/channels/qa-channel.md
|
||||
codeRefs:
|
||||
- extensions/qa-lab/src/report.ts
|
||||
- extensions/qa-lab/src/self-check.ts
|
||||
- src/agents/system-prompt.ts
|
||||
execution:
|
||||
kind: flow
|
||||
summary: Verify the agent can read repo docs and source, expand the QA plan, and publish a worked or did-not-work report.
|
||||
config:
|
||||
requiredFiles:
|
||||
- repo/qa/scenarios/index.md
|
||||
- repo/extensions/qa-lab/src/suite.ts
|
||||
- repo/docs/help/testing.md
|
||||
prompt: Read the seeded docs and source plan. The full repo is mounted under ./repo/. Explicitly inspect repo/qa/scenarios/index.md, repo/extensions/qa-lab/src/suite.ts, and repo/docs/help/testing.md, then report grouped into Worked, Failed, Blocked, and Follow-up. Mention at least two extra QA scenarios beyond the seed list.
|
||||
```
|
||||
|
||||
```yaml qa-flow
|
||||
steps:
|
||||
- name: reads seeded material and emits a protocol report
|
||||
actions:
|
||||
- call: reset
|
||||
- call: runAgentPrompt
|
||||
args:
|
||||
- ref: env
|
||||
- sessionKey: agent:qa:discovery
|
||||
message:
|
||||
expr: config.prompt
|
||||
timeoutMs:
|
||||
expr: liveTurnTimeoutMs(env, 30000)
|
||||
- call: waitForCondition
|
||||
saveAs: outbound
|
||||
args:
|
||||
- lambda:
|
||||
expr: "state.getSnapshot().messages.filter((candidate) => candidate.direction === 'outbound' && candidate.conversation.id === 'qa-operator' && hasDiscoveryLabels(candidate.text)).at(-1)"
|
||||
- expr: liveTurnTimeoutMs(env, 20000)
|
||||
- expr: "env.providerMode === 'mock-openai' ? 100 : 250"
|
||||
- assert:
|
||||
expr: "!reportsMissingDiscoveryFiles(outbound.text)"
|
||||
message:
|
||||
expr: "`discovery report still missed repo files: ${outbound.text}`"
|
||||
- assert:
|
||||
expr: "!reportsDiscoveryScopeLeak(outbound.text)"
|
||||
message:
|
||||
expr: "`discovery report drifted beyond scope: ${outbound.text}`"
|
||||
# Parity gate criterion 2 (no fake progress / fake tool completion):
|
||||
# require an actual read tool call before the prose report. Without this,
|
||||
# a model could fabricate a plausible Worked/Failed/Blocked/Follow-up
|
||||
# report without ever touching the repo files the prompt names. The
|
||||
# debug request log is fetched once and reused for both the assertion
|
||||
# and its failure-message diagnostic. Each request's allInputText is
|
||||
# lowercased inline at match time (the real prompt writes it as
|
||||
# "Worked, Failed, Blocked") so the contains check is case-insensitive.
|
||||
- set: discoveryDebugRequests
|
||||
value:
|
||||
expr: "env.mock ? [...(await fetchJson(`${env.mock.baseUrl}/debug/requests`))] : []"
|
||||
- assert:
|
||||
expr: "!env.mock || discoveryDebugRequests.some((request) => String(request.allInputText ?? '').toLowerCase().includes('worked, failed, blocked') && request.plannedToolName === 'read')"
|
||||
message:
|
||||
expr: "`expected at least one read tool call during discovery report scenario, saw plannedToolNames=${JSON.stringify(discoveryDebugRequests.map((request) => request.plannedToolName ?? null))}`"
|
||||
detailsExpr: outbound.text
|
||||
```
|
||||
Loading…
Add table
Add a link
Reference in a new issue