mirror of
https://github.com/hansjone/dsh-im-ops.git
synced 2026-10-09 00:33:20 +08:00
fix(im): fall back to workspace-file delivery when non-vision models reject images
This commit is contained in:
parent
e8c9817598
commit
0851e8cf30
6 changed files with 607 additions and 30 deletions
|
|
@ -6,6 +6,11 @@ This file records the notable changes in each dsh-im release. Its format follows
|
|||
|
||||
## [Unreleased]
|
||||
|
||||
### Added / 新增
|
||||
|
||||
- 非视觉模型收到图片时不再直接报错丢图:宿主以 `MODEL_DOES_NOT_SUPPORT_IMAGES` 拒绝带图片的 prompt 后,自动把同一批图片字节按入站文件管线落盘到 Session 工作区,并以"原文本 + 工具分析指引 + `<dsh_im_files>` 清单"的纯文本 prompt 复用同一 rpcId 重试一次,使非视觉模型仍可通过 run_code/pwsh 等工具识图;视觉模型与文件消息行为不变,其余图片错误仍按原样提示。
|
||||
Sending an image to a non-vision model no longer fails outright: when the Host rejects an image-bearing prompt with `MODEL_DOES_NOT_SUPPORT_IMAGES`, the same image bytes are automatically staged into the Session workspace through the inbound-file pipeline and retried once as a text-only prompt (original text plus tool-analysis guidance and the `<dsh_im_files>` manifest) under the same rpcId, so non-vision models can still inspect images via tools such as run_code/pwsh. Vision models and file messages are unchanged, and other image errors keep their existing messages.
|
||||
|
||||
## [4.4.0] - 2026-09-01
|
||||
|
||||
### Added / 新增
|
||||
|
|
|
|||
199
docs/方案/入站图片非视觉模型文件回退方案.md
Normal file
199
docs/方案/入站图片非视觉模型文件回退方案.md
Normal file
|
|
@ -0,0 +1,199 @@
|
|||
# 入站图片非视觉模型文件回退方案
|
||||
|
||||
> 状态:已实施(P0);Windows 本机全量回归通过,真机验收待各渠道补做
|
||||
>
|
||||
> 日期:2026-06-14
|
||||
>
|
||||
> 关联:[出站图片原生呈现落地方案](./出站图片原生呈现落地方案.md)、[渠道原生能力建设方案](./渠道原生能力建设方案.md)
|
||||
|
||||
## 1. 结论
|
||||
|
||||
IM 对话中,非视觉模型收到图片时当前是**准入即硬失败**:图片字节在宿主 `session.prompt` 准入检查处被整体丢弃,LLM 收不到任何内容;而发送文件(如 zip)却能成功。本方案不改宿主、不新建媒体框架,只做一件事:
|
||||
|
||||
**当宿主以 `MODEL_DOES_NOT_SUPPORT_IMAGES` 拒绝带图片的 prompt 时,dsh-im 自动把同一批图片字节转入现有入站文件管线(落盘到 Session 工作区 + `<dsh_im_files>` 清单),替换掉多模态内容块后重发一次 prompt。**
|
||||
|
||||
效果:
|
||||
|
||||
1. 非视觉模型不再对图片报错;图片以工作区文件形式到达 Agent,Agent 可用 `run_code`/`pwsh`(读字节、EXIF、OCR、图像库等)"以其他方式识图"。
|
||||
2. 视觉模型行为完全不变(继续原生多模态内容块)。
|
||||
3. 其他图片错误(超大、格式不支持、下载失败等)继续走现有报错,不回退。
|
||||
4. 收口在 `harness-client.mjs` 的 `ask()` 单点,九渠道 Bridge 无需各自改动。
|
||||
|
||||
与出站方案同样的原则:复用现有 `inbound-file` 生命周期(落盘、清单、turn 结束清理),不新增配置开关、不新增产物类型。
|
||||
|
||||
### 1.1 实施结果(2026-06-14)
|
||||
|
||||
本方案已按上述最小边界落地:
|
||||
|
||||
- `src/channels/shared/image-prompt.mjs` 新增 `IMAGE_FILE_FALLBACK_PROMPT`(模型侧指引,含英文翻译)、`imageFileSourcesFromContent()`(图片内容块 → 文件源,含扩展名映射与文件名清洗)、`contentWithoutImages()`、`isModelImageRejection()`。
|
||||
- `src/channels/shared/harness-client.mjs` 的 `ask()` 将原入站文件落盘逻辑抽为 `#stageWorkspaceFiles()`;`session.prompt` 被以 `MODEL_DOES_NOT_SUPPORT_IMAGES` 拒绝且 content 含图片块时,把同一批字节经该管线落盘、重建纯文本 prompt(原文本 + 指引 + 合并后的单个 `<dsh_im_files>` 清单)并复用同一 `promptRpcId` 重试一次;重试不再回退,落盘失败时保留原始错误文案。staged 批次统一进入既有 turn 结束清理。
|
||||
- `test/image-fallback.test.mjs` 新增 8 个用例:转换/判定辅助函数、完整回退(复用 rpcId、字节一致、清单正确、turn 后清理)、混合消息合并清单、非模型原因不回退、落盘失败保留原错误、重试失败不再重试、纯文本路径不变。
|
||||
- 回归:Windows 本机以逐文件直跑方式执行全部 137 个测试文件,失败集与改动前基线(git worktree 对照)完全一致(23 个均为仓库既有的 Windows/沙箱环境性失败,CI 在 ubuntu 上通过);`npm run build` 通过;`verify-package` 的可执行位检查在 Windows 上为既有环境性失败。
|
||||
|
||||
## 2. 现状与根因
|
||||
|
||||
### 2.1 入站文件链路(zip 能成功)
|
||||
|
||||
```
|
||||
message.files ──► harness.ask(…, { files })
|
||||
└► fileIngressExecutor(宿主进程内执行)
|
||||
└► stageInboundFiles() 写入 <workspace>/.dsh-im/inbound/turn-XXXX/NN-名字(0600)
|
||||
└► appendInboundFilesToPrompt() 在 prompt 末尾追加 <dsh_im_files> JSON 清单(工作区相对路径)
|
||||
└► session.prompt(纯文本内容块)
|
||||
```
|
||||
|
||||
- 代码:`src/channels/shared/inbound-file.mjs`(`stageInboundFiles`、`appendInboundFilesToPrompt`)、`src/channels/shared/harness-client.mjs` ask() 内 `inboundFiles` 分支、`plugin-src/host/harness-session-coordinator.mjs` 的 `createFileIngressExecutor`。
|
||||
- 纯文本 prompt 不依赖模型输入模态,任何模型都能收到路径清单,再用工具处理文件。
|
||||
|
||||
### 2.2 入站图片链路(被拦截)
|
||||
|
||||
```
|
||||
message.images ──► promptContentForMessage()(Bridge 内,含大小/数量/魔数校验)
|
||||
└► content = [{type:'text'}, {type:'image', mediaType, data: base64}]
|
||||
└► session.prompt RPC
|
||||
└► 宿主 dsh-host-apiproxy prompt 准入:
|
||||
resolveModelInfo(provider, model).inputModalities 不含 'image'
|
||||
──► 返回 attachment-error / MODEL_DOES_NOT_SUPPORT_IMAGES
|
||||
(发生在 createUserMessage 之前,未产生任何 durable 消息)
|
||||
└► HarnessRpcError 冒泡回 Bridge
|
||||
└► imagePromptUserMessage() 翻译为用户文案:
|
||||
"当前模型不支持图片,请用 /models 查看可用模型,再用 /model <序号> 切换后重发。"
|
||||
```
|
||||
|
||||
- 代码:`src/channels/shared/image-prompt.mjs`(`promptContentForMessage`、`HOST_ATTACHMENT_USER_MESSAGES`);宿主侧为 `@deepseek-ai/dsh-host-apiproxy` prompt handler(`hasImage` → `resolveModelInfo` 检查 → `err`)。
|
||||
- 六个图片构建点(共享 `text-harness-bridge.mjs` 及飞书/QQ/钉钉/微信/企业微信各自 Bridge)全部经 `harness.ask()` 收口,因此回退只需改一处。
|
||||
|
||||
### 2.3 不对称的根源
|
||||
|
||||
| 维度 | 文件(zip) | 图片 |
|
||||
| --- | --- | --- |
|
||||
| 交付形态 | 工作区路径文本 | 多模态内容块 |
|
||||
| 对模型能力的要求 | 无 | `inputModalities` 含 `image` |
|
||||
| 失败时的降级 | 不适用 | **无**(直接丢弃字节) |
|
||||
|
||||
另一个限制:`session.models` / `llm.models` 返回的模型目录只有 `id/name/description/reasoning`,**不含输入模态**,dsh-im 无法预判当前模型是否支持图片;唯一的信号就是宿主拒绝时的 `MODEL_DOES_NOT_SUPPORT_IMAGES` 错误码——这也正是回退的触发器。
|
||||
|
||||
## 3. 设计原则
|
||||
|
||||
1. **图片字节必须到达 Session 工作区**,而不是被丢弃——这是"非视觉模型以其他方式识图"的前提。
|
||||
2. **不修改宿主**。`@deepseek-ai/dsh-host-apiproxy` 是安装的 npm 包,改了会被升级覆盖;宿主侧改进作为上游建议另行反馈(见 §7)。
|
||||
3. **复用入站文件管线**,不创建第二条图片生命周期:同样的目录结构、权限(0600)、turn 结束清理语义。
|
||||
4. **单点收口**:回退逻辑只存在于 `harness-client.mjs` `ask()` 内,所有渠道自动获益。
|
||||
5. **不新增配置开关**:对 `MODEL_DOES_NOT_SUPPORT_IMAGES` 的回退严格优于硬失败,默认启用;其余错误原因维持现有行为。
|
||||
6. **至多重试一次**,防止错误码异常时的循环。
|
||||
|
||||
## 4. 方案设计
|
||||
|
||||
### 4.1 触发条件
|
||||
|
||||
在 `ask()` 中包装 `session.prompt` 调用,捕获 `HarnessRpcError`,当且仅当:
|
||||
|
||||
- `error.code === 'attachment-error'` **且**
|
||||
- `error.details.reason === 'MODEL_DOES_NOT_SUPPORT_IMAGES'` **且**
|
||||
- 本次 `content` 含 `{type:'image'}` 内容块
|
||||
|
||||
进入回退;否则原样抛出。
|
||||
|
||||
### 4.2 图片 → 文件源转换
|
||||
|
||||
从被拒 `content` 的 image 内容块直接构造文件源(字节已在内存,**无第二次网络下载**):
|
||||
|
||||
```js
|
||||
// image-prompt.mjs 新增 helper(图片知识内聚;stageInboundFiles 无需改动)
|
||||
const IMAGE_EXTENSIONS = new Map([
|
||||
['image/png', '.png'],
|
||||
['image/jpeg', '.jpg'],
|
||||
['image/gif', '.gif'],
|
||||
['image/webp', '.webp'],
|
||||
]);
|
||||
|
||||
export function imageFileSourcesFromContent(content) {
|
||||
return content
|
||||
.filter((part) => part?.type === 'image')
|
||||
.map((part, index) => ({
|
||||
name: part.name ?? ('image-' + (index + 1) + (IMAGE_EXTENSIONS.get(part.mediaType) ?? '.img')),
|
||||
mediaType: part.mediaType,
|
||||
data: Buffer.from(part.data, 'base64'),
|
||||
}));
|
||||
}
|
||||
```
|
||||
|
||||
- `stageInboundFiles` 的 `loadedFile()` 已接受 `{data: Buffer, name, mediaType}` 形态,无需改动即可落盘。
|
||||
- `storageName()` 已做文件名清洗;扩展名映射保证下游工具能按扩展名识别格式。
|
||||
- 大小限制天然满足:`promptContentForMessage` 已按单张 5MB / 总量 20MB / 20 张校验过,转换不放大。
|
||||
|
||||
### 4.3 重发 prompt
|
||||
|
||||
1. 保留 append 文件清单**之前**的 `basePrompt`(当前代码在 ask() 内先 `prompt = appendInboundFilesToPrompt(prompt, stagedInboundFiles)` 再发 prompt;回退需从原始 prompt 重建,避免出现两个 `<dsh_im_files>` 块)。
|
||||
2. 新 `content` = 原 text 内容块 + 回退提示文本 + 合并后的单个 `<dsh_im_files>` 清单(原 staged 文件 + 图片文件,一次 `appendInboundFilesToPrompt` 生成)。
|
||||
3. 提示文本(走 `t()`,中英双语):
|
||||
|
||||
```
|
||||
当前会话模型不支持直接接收图片输入。用户发送的图片已作为文件保存到工作区(见下方清单)。
|
||||
请使用可用工具分析这些图片文件后回答,例如 run_code/pwsh 读取字节、解析元数据、
|
||||
调用图像处理或 OCR 库;不要假设自己能直接看到图片内容。
|
||||
```
|
||||
|
||||
4. 再次调用 `session.prompt`(此时内容为纯文本,宿主不再走图片准入分支,也不受 `serializeImageAdmission` 串行化影响)。
|
||||
|
||||
### 4.4 生命周期与安全
|
||||
|
||||
- **幂等**:宿主检查发生在 `createUserMessage` 之前,被拒的 prompt 未入队、未产生 durable 消息,重发不会造成重复消息或空 turn。
|
||||
- **清理**:图片文件与现有入站文件同语义——turn 结束后由既有 `stagedInboundFiles.cleanup()` 一并删除(回退实现需把图片 staged 结果并入同一清理集合;重试仍失败时立即清理)。
|
||||
- **降级失败**:若 `#fileIngressExecutor` 不可用(`inbound-file-ingress-unavailable`)或落盘失败,则不重试,按现有链路向用户报错(保留 `MODEL_DOES_NOT_SUPPORT_IMAGES` 的既有文案作为兜底,可追加"图片转文件失败"说明)。
|
||||
- **旧宿主兼容**:错误码不存在时(旧版本宿主不检查模态或行为不同)回退不触发,行为与今天一致。
|
||||
|
||||
### 4.5 用户可见行为
|
||||
|
||||
- 回退成功后**不再报错**,Agent 正常处理并回复;可选(P1):通过现有 `onUpdate` 流在开头推一条一次性提示"当前模型不支持直接识图,已将图片保存为工作区文件进行工具分析",让用户知情。
|
||||
- `/models` 文案与 `HOST_ATTACHMENT_USER_MESSAGES` 映射保留:仅在回退自身失败时作为兜底呈现。
|
||||
|
||||
### 4.6 改动清单
|
||||
|
||||
| 文件 | 改动 | 量级 |
|
||||
| --- | --- | --- |
|
||||
| `src/channels/shared/harness-client.mjs` | ask() 内 `session.prompt` 包回退:触发判断、保留 basePrompt、图片落盘、清单合并、重试一次、清理集合 | ~80 行 |
|
||||
| `src/channels/shared/image-prompt.mjs` | 新增 `imageFileSourcesFromContent()`、扩展名映射 | ~30 行 |
|
||||
| `src/channels/shared/inbound-file.mjs` | (若需要)导出合并两个 staged 清单的 helper | ~10 行 |
|
||||
| `src/channels/shared/i18n-en/*.mjs` | 回退提示文本英文翻译 | 若干 |
|
||||
| `test/image-fallback.test.mjs`(新增)、`test/inbound-file.test.mjs` | 见 §6 | ~150 行 |
|
||||
|
||||
## 5. 备选方案对比
|
||||
|
||||
| 方案 | 说明 | 取舍 |
|
||||
| --- | --- | --- |
|
||||
| **A. 客户端响应式回退(本方案)** | 收到 `MODEL_DOES_NOT_SUPPORT_IMAGES` 后转文件重发 | ✅ 不改宿主、单点收口、旧宿主兼容;代价是多一次被拒往返 |
|
||||
| B. 宿主准入降级 | apiproxy 在准入处把 image part 换成附件路径文本而非拒绝 | 效果最好(所有 API 客户端获益),但改的是安装的 npm 包,升级即丢;建议作为上游 issue 反馈,dsh-im 不依赖它 |
|
||||
| C. 模型目录暴露 `inputModalities` | 宿主在 `session.models`/`llm.models` 增加模态字段 | 可让 dsh-im 预判、直接走文件路径省一次往返,`/models` 可加视觉标记;需宿主支持,旧宿主下仍靠 A 兜底;作为上游建议一并反馈 |
|
||||
| D. 纯文案优化 | 不回退,只改报错文案教用户切换模型 | 不解决"非视觉模型识图可能性"被阻断的问题 |
|
||||
|
||||
推荐:**本期实施 A**;B、C 作为上游反馈(对 DSH 宿主仓库提 issue),A 在 C 落地后仍作为旧宿主兜底保留。
|
||||
|
||||
## 6. 测试与验收
|
||||
|
||||
### 6.1 单元测试(新增 `test/image-fallback.test.mjs`,fake rpc/executor)
|
||||
|
||||
1. 首次 `session.prompt` 返回 `attachment-error/MODEL_DOES_NOT_SUPPORT_IMAGES` → 断言:调用了第二次 prompt;第二次 content 无 image part;含 `<dsh_im_files>` 清单与回退提示文本;文件以正确扩展名写入 executor 收到的 sources;仅重试一次。
|
||||
2. 混合消息(图片+文件)→ 断言只出现**一个**合并后的 `<dsh_im_files>` 块。
|
||||
3. 其他拒绝原因(`IMAGE_TOO_LARGE`、`TOO_MANY_IMAGES` 等)→ 不回退,原错误冒泡,用户文案不变。
|
||||
4. 重试再失败(如 executor 抛 `inbound-file-ingress-unavailable`)→ 报错且 staged 文件被清理。
|
||||
5. 纯文本/纯文件消息 → 不触发任何回退路径(回归)。
|
||||
6. turn 正常结束后 staged 图片文件被 cleanup(复用 `test/inbound-file.test.mjs` 的断言模式)。
|
||||
|
||||
### 6.2 回归与手工验收
|
||||
|
||||
- `npm run check` 全量通过。
|
||||
- 真机:绑定非视觉模型的会话发一张图 → 收到 Agent 基于工具分析的回复而非报错;视觉模型会话发图 → 行为与现状一致(原生多模态);发 zip → 行为不变。
|
||||
|
||||
## 7. 上游反馈(非本仓库范围)
|
||||
|
||||
向 DSH 宿主(`@deepseek-ai/dsh-host-apiproxy`)反馈两个改进建议:
|
||||
|
||||
1. prompt 准入遇到非视觉模型时,将 image part 降级为持久化附件路径文本(对应方案 B),使所有 IM/API 客户端天然获得该能力;
|
||||
2. 模型目录响应增加 `inputModalities`(对应方案 C),让客户端可预判并在 `/models` 中展示视觉能力标记。
|
||||
|
||||
## 8. 分期
|
||||
|
||||
- **P0(本期)**:方案 A 全部内容 + 单测 + 真机验收。
|
||||
- **P1**:回退触发时的用户可见一次性提示;`<dsh_im_files>` 清单 description 中补充"非视觉会话请用工具分析图片"的指引(依赖 C 落地后可精化为按模型分流)。
|
||||
- **P2**:上游 B/C 跟进结果回流后,评估是否保留响应式回退作为兜底(预期保留)。
|
||||
|
|
@ -7,6 +7,12 @@ import {
|
|||
appendInboundFilesToPrompt,
|
||||
InboundFileError,
|
||||
} from './inbound-file.mjs';
|
||||
import {
|
||||
IMAGE_FILE_FALLBACK_PROMPT,
|
||||
contentWithoutImages,
|
||||
imageFileSourcesFromContent,
|
||||
isModelImageRejection,
|
||||
} from './image-prompt.mjs';
|
||||
import { outboundArtifactRegistry } from './semantic/artifact.mjs';
|
||||
import { t } from './i18n.mjs';
|
||||
import { watchHarnessMux } from './harness-mux.mjs';
|
||||
|
|
@ -1236,6 +1242,31 @@ export class HarnessClient {
|
|||
return ownership ? { ownership, recovered: true } : null;
|
||||
}
|
||||
|
||||
/** Stage inbound file sources into the Session workspace via the Host executor. */
|
||||
async #stageWorkspaceFiles(sessionId, files, signal) {
|
||||
if (!this.#fileIngressExecutor) {
|
||||
throw new InboundFileError(
|
||||
'inbound-file-ingress-unavailable',
|
||||
'Harness file ingress is unavailable in this Host process.',
|
||||
);
|
||||
}
|
||||
const sessionList = await this.rpc(
|
||||
'session.list',
|
||||
{},
|
||||
30_000,
|
||||
{ signal },
|
||||
);
|
||||
const sessionWorkspace = sessionList?.items?.find(
|
||||
(item) => item?.sessionId === sessionId,
|
||||
)?.cwd;
|
||||
return this.#fileIngressExecutor({
|
||||
sessionId,
|
||||
workspace: sessionWorkspace,
|
||||
files,
|
||||
signal,
|
||||
});
|
||||
}
|
||||
|
||||
async ask(sessionId, prompt, options = {}) {
|
||||
if (typeof options === 'number') options = { timeoutMs: options };
|
||||
const timeoutMs = options.timeoutMs ?? 600_000;
|
||||
|
|
@ -1292,7 +1323,7 @@ export class HarnessClient {
|
|||
let interactionTask = null;
|
||||
let artifactsDelivered = false;
|
||||
let deliveredArtifactCount = 0;
|
||||
let stagedInboundFiles = null;
|
||||
const stagedBatches = [];
|
||||
let promptAccepted = false;
|
||||
let turnFinished = false;
|
||||
|
||||
|
|
@ -1323,29 +1354,11 @@ export class HarnessClient {
|
|||
const closeArtifactConsumer = outboundArtifactRegistry.openConsumer(sessionId, promptRpcId);
|
||||
|
||||
try {
|
||||
const basePrompt = prompt;
|
||||
if (inboundFiles.length > 0) {
|
||||
if (!this.#fileIngressExecutor) {
|
||||
throw new InboundFileError(
|
||||
'inbound-file-ingress-unavailable',
|
||||
'Harness file ingress is unavailable in this Host process.',
|
||||
);
|
||||
}
|
||||
const sessionList = await this.rpc(
|
||||
'session.list',
|
||||
{},
|
||||
30_000,
|
||||
{ signal },
|
||||
);
|
||||
const sessionWorkspace = sessionList?.items?.find(
|
||||
(item) => item?.sessionId === sessionId,
|
||||
)?.cwd;
|
||||
stagedInboundFiles = await this.#fileIngressExecutor({
|
||||
sessionId,
|
||||
workspace: sessionWorkspace,
|
||||
files: inboundFiles,
|
||||
signal,
|
||||
});
|
||||
prompt = appendInboundFilesToPrompt(prompt, stagedInboundFiles);
|
||||
const staged = await this.#stageWorkspaceFiles(sessionId, inboundFiles, signal);
|
||||
stagedBatches.push(staged);
|
||||
prompt = appendInboundFilesToPrompt(prompt, staged);
|
||||
}
|
||||
if (interactionSignal) {
|
||||
let markOpen;
|
||||
|
|
@ -1372,12 +1385,47 @@ export class HarnessClient {
|
|||
if (!Array.isArray(content) || content.length === 0) {
|
||||
throw new TypeError('Harness prompt content is required');
|
||||
}
|
||||
await this.rpc('session.prompt', {
|
||||
const clientTimeZone = Intl.DateTimeFormat().resolvedOptions().timeZone;
|
||||
const sendPrompt = (promptContent) => this.rpc('session.prompt', {
|
||||
sessionId,
|
||||
mode: 'queue',
|
||||
content,
|
||||
clientTimeZone: Intl.DateTimeFormat().resolvedOptions().timeZone,
|
||||
content: promptContent,
|
||||
clientTimeZone,
|
||||
}, 30_000, { rpcId: promptRpcId, signal });
|
||||
try {
|
||||
await sendPrompt(content);
|
||||
} catch (error) {
|
||||
// The Host refuses image blocks for a non-vision model before any
|
||||
// durable user message exists. Re-deliver the same bytes the way
|
||||
// ordinary uploads (zip, documents) already travel — staged into the
|
||||
// Session workspace and named in a text manifest — then retry once
|
||||
// with a text-only prompt. The retry reuses promptRpcId so reply
|
||||
// tracking, control and interaction ownership stay bound to this ask.
|
||||
const imageSources = isModelImageRejection(error)
|
||||
? imageFileSourcesFromContent(content)
|
||||
: [];
|
||||
if (imageSources.length === 0) throw error;
|
||||
let stagedImages;
|
||||
try {
|
||||
stagedImages = await this.#stageWorkspaceFiles(sessionId, imageSources, signal);
|
||||
} catch (stagingError) {
|
||||
if (signal?.aborted) throw error;
|
||||
console.warn(
|
||||
`[${this.#logPrefix}] unable to restage rejected images as workspace files:`,
|
||||
stagingError?.message ?? String(stagingError),
|
||||
);
|
||||
throw error;
|
||||
}
|
||||
stagedBatches.push(stagedImages);
|
||||
const baseContent = typeof basePrompt === 'string'
|
||||
? [{ type: 'text', text: basePrompt }]
|
||||
: basePrompt;
|
||||
const fallbackPrompt = appendInboundFilesToPrompt([
|
||||
...contentWithoutImages(baseContent),
|
||||
{ type: 'text', text: t(IMAGE_FILE_FALLBACK_PROMPT) },
|
||||
], { files: stagedBatches.flatMap((batch) => batch?.files ?? []) });
|
||||
await sendPrompt(fallbackPrompt);
|
||||
}
|
||||
promptAccepted = true;
|
||||
|
||||
try {
|
||||
|
|
@ -1435,10 +1483,14 @@ export class HarnessClient {
|
|||
throw turnStoppedError();
|
||||
}
|
||||
} finally {
|
||||
if (stagedInboundFiles && (!promptAccepted || turnFinished)) {
|
||||
await stagedInboundFiles.cleanup().catch((error) => {
|
||||
console.warn(`[${this.#logPrefix}] unable to clean inbound files:`, error.message);
|
||||
});
|
||||
if (!promptAccepted || turnFinished) {
|
||||
for (const staged of stagedBatches) {
|
||||
try {
|
||||
await staged?.cleanup?.();
|
||||
} catch (error) {
|
||||
console.warn(`[${this.#logPrefix}] unable to clean inbound files:`, error.message);
|
||||
}
|
||||
}
|
||||
}
|
||||
closeArtifactConsumer();
|
||||
if (ownership) {
|
||||
|
|
|
|||
|
|
@ -81,6 +81,8 @@ export default {
|
|||
// image-prompt.mjs
|
||||
'当前模型不支持图片,请用 /models 查看可用模型,再用 /model <序号> 切换后重发。':
|
||||
'The current model does not support images. Use /models to list available models, switch with /model <number>, then resend.',
|
||||
'当前会话模型不支持直接接收图片输入。用户发送的图片已作为文件保存到工作区(见下方文件清单)。请使用可用工具分析这些图片文件后回答,例如 run_code 或 pwsh 读取字节、解析元数据、调用图像处理或 OCR 库;不要假设自己能直接看到图片内容。':
|
||||
'The current session model does not accept direct image input. The images sent by the user were saved into the workspace as files (see the file manifest below). Answer by analyzing those image files with the available tools — for example run_code or pwsh to read bytes, parse metadata, or call image-processing or OCR libraries — and do not assume you can see the images directly.',
|
||||
'图片超过宿主允许的大小,请压缩后重试。':
|
||||
'The image exceeds the size allowed by the host; compress it and try again.',
|
||||
'图片分辨率过高,请压缩后重试。':
|
||||
|
|
|
|||
|
|
@ -6,6 +6,12 @@ const DEFAULT_MAX_TOTAL_IMAGE_BYTES = 20 * 1024 * 1024;
|
|||
|
||||
export const DEFAULT_IMAGE_PROMPT = '请分析这张图片。';
|
||||
|
||||
/**
|
||||
* Model-facing guidance appended when the Host refuses image input for the
|
||||
* current model and the same images are re-delivered as workspace files.
|
||||
*/
|
||||
export const IMAGE_FILE_FALLBACK_PROMPT = '当前会话模型不支持直接接收图片输入。用户发送的图片已作为文件保存到工作区(见下方文件清单)。请使用可用工具分析这些图片文件后回答,例如 run_code 或 pwsh 读取字节、解析元数据、调用图像处理或 OCR 库;不要假设自己能直接看到图片内容。';
|
||||
|
||||
export class ImagePromptError extends Error {
|
||||
constructor(code, message, userMessage, options = {}) {
|
||||
super(message, options);
|
||||
|
|
@ -299,3 +305,48 @@ export function imagePromptDiagnostic(error) {
|
|||
export function imagePromptUserMessage(error) {
|
||||
return imagePromptDiagnostic(error)?.userMessage ?? null;
|
||||
}
|
||||
|
||||
const IMAGE_FILE_EXTENSIONS = new Map([
|
||||
['image/png', '.png'],
|
||||
['image/jpeg', '.jpg'],
|
||||
['image/gif', '.gif'],
|
||||
['image/webp', '.webp'],
|
||||
]);
|
||||
|
||||
const IMAGE_EXTENSION_PATTERN = /\.(?:png|jpe?g|gif|webp)$/i;
|
||||
|
||||
function imageStorageName(name, mediaType, index) {
|
||||
const extension = IMAGE_FILE_EXTENSIONS.get(mediaType) ?? '.img';
|
||||
const cleaned = safeName(name);
|
||||
if (cleaned && IMAGE_EXTENSION_PATTERN.test(cleaned)) return cleaned;
|
||||
return `${cleaned ?? `image-${index + 1}`}${extension}`;
|
||||
}
|
||||
|
||||
/**
|
||||
* Convert already-admitted image content blocks into inbound file sources so
|
||||
* the same bytes can reach a non-vision model as workspace files — the path
|
||||
* ordinary uploads such as zip archives already take.
|
||||
*/
|
||||
export function imageFileSourcesFromContent(content) {
|
||||
if (!Array.isArray(content)) return [];
|
||||
return content
|
||||
.filter((part) => part?.type === 'image')
|
||||
.map((part, index) => ({
|
||||
name: imageStorageName(part.name, part.mediaType, index),
|
||||
...(typeof part.mediaType === 'string' && part.mediaType.trim()
|
||||
? { mediaType: part.mediaType.trim() }
|
||||
: {}),
|
||||
data: Buffer.from(typeof part.data === 'string' ? part.data : '', 'base64'),
|
||||
}));
|
||||
}
|
||||
|
||||
/** Return the same content with every image block removed. */
|
||||
export function contentWithoutImages(content) {
|
||||
return Array.isArray(content) ? content.filter((part) => part?.type !== 'image') : content;
|
||||
}
|
||||
|
||||
/** Whether an error is the Host rejecting image input for a non-vision model. */
|
||||
export function isModelImageRejection(error) {
|
||||
return error?.code === 'attachment-error'
|
||||
&& error?.details?.reason === 'MODEL_DOES_NOT_SUPPORT_IMAGES';
|
||||
}
|
||||
|
|
|
|||
268
test/image-fallback.test.mjs
Normal file
268
test/image-fallback.test.mjs
Normal file
|
|
@ -0,0 +1,268 @@
|
|||
import assert from 'node:assert/strict';
|
||||
import {
|
||||
mkdtemp,
|
||||
readFile,
|
||||
rm,
|
||||
} from 'node:fs/promises';
|
||||
import { tmpdir } from 'node:os';
|
||||
import { join, resolve } from 'node:path';
|
||||
import test from 'node:test';
|
||||
|
||||
import {
|
||||
IMAGE_FILE_FALLBACK_PROMPT,
|
||||
contentWithoutImages,
|
||||
imageFileSourcesFromContent,
|
||||
isModelImageRejection,
|
||||
} from '../src/channels/shared/image-prompt.mjs';
|
||||
import { stageInboundFiles } from '../src/channels/shared/inbound-file.mjs';
|
||||
import {
|
||||
HarnessClient,
|
||||
HarnessRpcError,
|
||||
} from '../src/channels/shared/harness-client.mjs';
|
||||
|
||||
const PNG_BYTES = Buffer.from(
|
||||
'iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAQAAAC1HAwCAAAAC0lEQVR42mNk+A8AAQUBAScY42YAAAAASUVORK5CYII=',
|
||||
'base64',
|
||||
);
|
||||
|
||||
function imageContent() {
|
||||
return [
|
||||
{ type: 'text', text: 'what is in this picture?' },
|
||||
{ type: 'image', mediaType: 'image/png', data: PNG_BYTES.toString('base64'), name: 'photo.png' },
|
||||
];
|
||||
}
|
||||
|
||||
function modelImageRejection() {
|
||||
return new HarnessRpcError('session.prompt', {
|
||||
code: 'attachment-error',
|
||||
message: 'Model "text-only" does not support image input.',
|
||||
details: { reason: 'MODEL_DOES_NOT_SUPPORT_IMAGES' },
|
||||
});
|
||||
}
|
||||
|
||||
async function workspace(t) {
|
||||
const directory = await mkdtemp(join(tmpdir(), 'dsh-im-image-fallback-'));
|
||||
t.after(() => rm(directory, { recursive: true, force: true }));
|
||||
return directory;
|
||||
}
|
||||
|
||||
/** A client whose RPC layer is scripted per test; staging uses the real pipeline. */
|
||||
function scriptedClient({ workspaceRoot, fileIngressExecutor, onPrompt }) {
|
||||
const client = new HarnessClient({
|
||||
baseUrl: 'http://127.0.0.1:3080',
|
||||
workspace: workspaceRoot,
|
||||
...(fileIngressExecutor !== undefined ? { fileIngressExecutor } : {}),
|
||||
});
|
||||
client.ensureRunning = async () => true;
|
||||
const promptCalls = [];
|
||||
let historyCalls = 0;
|
||||
client.rpc = async (method, payload, _timeoutMs, options) => {
|
||||
if (method === 'session.prompt') {
|
||||
promptCalls.push({ payload, rpcId: options.rpcId });
|
||||
return onPrompt(promptCalls.length, options.rpcId);
|
||||
}
|
||||
if (method === 'session.list') {
|
||||
return { items: [{ sessionId: payload.sessionId ?? 'session-fallback', cwd: workspaceRoot }] };
|
||||
}
|
||||
assert.equal(method, 'session.history', `unexpected RPC: ${method}`);
|
||||
historyCalls += 1;
|
||||
if (historyCalls === 1) return { events: [] };
|
||||
return {
|
||||
events: [
|
||||
{ event: { seq: 1, type: 'turn/start', data: { turn: 7 } } },
|
||||
{ event: {
|
||||
seq: 2,
|
||||
type: 'user/message',
|
||||
data: { turn: 7, source: { rpcId: promptCalls[0]?.rpcId ?? options.rpcId } },
|
||||
} },
|
||||
{ event: {
|
||||
seq: 3,
|
||||
type: 'assistant/message',
|
||||
data: { turn: 7, message: { content: [{ type: 'text', text: 'answer from tools' }] } },
|
||||
} },
|
||||
{ event: { seq: 4, type: 'turn/end', data: { turn: 7, reason: 'completed' } } },
|
||||
],
|
||||
};
|
||||
};
|
||||
return { client, promptCalls };
|
||||
}
|
||||
|
||||
function manifestFiles(content) {
|
||||
const parts = content
|
||||
.filter((part) => part?.type === 'text' && part.text.includes('<dsh_im_files>'));
|
||||
assert.equal(parts.length, 1, 'exactly one <dsh_im_files> manifest block');
|
||||
return JSON.parse(parts[0].text.split('\n').find((line) => line.startsWith('{'))).files;
|
||||
}
|
||||
|
||||
test('imageFileSourcesFromContent maps image blocks to safe file sources', () => {
|
||||
assert.deepEqual(imageFileSourcesFromContent([
|
||||
{ type: 'text', text: 'caption' },
|
||||
{ type: 'image', mediaType: 'image/png', data: PNG_BYTES.toString('base64'), name: '../photo.png' },
|
||||
{ type: 'image', mediaType: 'image/jpeg', data: '' },
|
||||
{ type: 'image', mediaType: 'image/webp', data: '', name: 'shot' },
|
||||
{ type: 'image', mediaType: 'image/gif', data: '', name: 'anim.gif' },
|
||||
null,
|
||||
]), [
|
||||
{ name: 'photo.png', mediaType: 'image/png', data: PNG_BYTES },
|
||||
{ name: 'image-2.jpg', mediaType: 'image/jpeg', data: Buffer.alloc(0) },
|
||||
{ name: 'shot.webp', mediaType: 'image/webp', data: Buffer.alloc(0) },
|
||||
{ name: 'anim.gif', mediaType: 'image/gif', data: Buffer.alloc(0) },
|
||||
]);
|
||||
|
||||
assert.deepEqual(contentWithoutImages(imageContent()), [
|
||||
{ type: 'text', text: 'what is in this picture?' },
|
||||
]);
|
||||
assert.deepEqual(imageFileSourcesFromContent('plain text'), []);
|
||||
assert.equal(contentWithoutImages('plain text'), 'plain text');
|
||||
});
|
||||
|
||||
test('isModelImageRejection matches only the non-vision admission reason', () => {
|
||||
assert.equal(isModelImageRejection(modelImageRejection()), true);
|
||||
assert.equal(isModelImageRejection({
|
||||
code: 'attachment-error',
|
||||
details: { reason: 'IMAGE_TOO_LARGE' },
|
||||
}), false);
|
||||
assert.equal(isModelImageRejection({
|
||||
code: 'agent-busy',
|
||||
details: { reason: 'MODEL_DOES_NOT_SUPPORT_IMAGES' },
|
||||
}), false);
|
||||
assert.equal(isModelImageRejection(new Error('unrelated')), false);
|
||||
});
|
||||
|
||||
test('HarnessClient restages rejected images as workspace files and retries text-only', async (t) => {
|
||||
const root = await workspace(t);
|
||||
const ingressCalls = [];
|
||||
const { client, promptCalls } = scriptedClient({
|
||||
workspaceRoot: root,
|
||||
fileIngressExecutor: ({ files, workspace, signal }) => {
|
||||
ingressCalls.push({ files, workspace });
|
||||
return stageInboundFiles({ files }, { workspace, signal });
|
||||
},
|
||||
onPrompt: (attempt) => (attempt === 1 ? Promise.reject(modelImageRejection()) : Promise.resolve({})),
|
||||
});
|
||||
|
||||
assert.equal(
|
||||
await client.ask('session-fallback', imageContent(), { timeoutMs: 3_000 }),
|
||||
'answer from tools',
|
||||
);
|
||||
|
||||
// Exactly one retry, reusing the same rpcId so reply tracking stays bound.
|
||||
assert.equal(promptCalls.length, 2);
|
||||
assert.equal(promptCalls[1].rpcId, promptCalls[0].rpcId);
|
||||
assert.deepEqual(promptCalls[0].payload.content, imageContent());
|
||||
|
||||
const retryContent = promptCalls[1].payload.content;
|
||||
assert.equal(retryContent.some((part) => part?.type === 'image'), false);
|
||||
assert.deepEqual(retryContent[0], { type: 'text', text: 'what is in this picture?' });
|
||||
assert.deepEqual(retryContent[1], { type: 'text', text: IMAGE_FILE_FALLBACK_PROMPT });
|
||||
|
||||
const files = manifestFiles(retryContent);
|
||||
assert.deepEqual(files.map(({ name }) => name), ['photo.png']);
|
||||
assert.match(files[0].path, /^\.dsh-im[\\/]inbound[\\/]turn-[^\\/]+[\\/]01-photo\.png$/);
|
||||
|
||||
// The exact image bytes were staged into the Session workspace via ingress.
|
||||
assert.equal(ingressCalls.length, 1);
|
||||
assert.equal(ingressCalls[0].workspace, root);
|
||||
assert.equal(ingressCalls[0].files[0].data.equals(PNG_BYTES), true);
|
||||
|
||||
// The turn finishing cleans the staged image like any inbound file.
|
||||
await assert.rejects(readFile(resolve(root, files[0].path)), /ENOENT/);
|
||||
});
|
||||
|
||||
test('mixed upload and image messages merge into one manifest after fallback', async (t) => {
|
||||
const root = await workspace(t);
|
||||
const { client, promptCalls } = scriptedClient({
|
||||
workspaceRoot: root,
|
||||
fileIngressExecutor: ({ files, workspace, signal }) => (
|
||||
stageInboundFiles({ files }, { workspace, signal })
|
||||
),
|
||||
onPrompt: (attempt) => (attempt === 1 ? Promise.reject(modelImageRejection()) : Promise.resolve({})),
|
||||
});
|
||||
|
||||
assert.equal(await client.ask('session-fallback', imageContent(), {
|
||||
timeoutMs: 3_000,
|
||||
files: [{ name: 'archive.zip', data: Buffer.from('PK\u0003\u0004'), mediaType: 'application/zip' }],
|
||||
}), 'answer from tools');
|
||||
|
||||
assert.equal(promptCalls.length, 2);
|
||||
const retryContent = promptCalls[1].payload.content;
|
||||
assert.equal(retryContent.some((part) => part?.type === 'image'), false);
|
||||
assert.deepEqual(manifestFiles(retryContent).map(({ name }) => name), [
|
||||
'archive.zip',
|
||||
'photo.png',
|
||||
]);
|
||||
});
|
||||
|
||||
test('other image admission failures never trigger the fallback', async (t) => {
|
||||
const root = await workspace(t);
|
||||
let ingressUsed = false;
|
||||
const { client, promptCalls } = scriptedClient({
|
||||
workspaceRoot: root,
|
||||
fileIngressExecutor: async () => {
|
||||
ingressUsed = true;
|
||||
throw new Error('must not stage');
|
||||
},
|
||||
onPrompt: () => Promise.reject(new HarnessRpcError('session.prompt', {
|
||||
code: 'attachment-error',
|
||||
message: 'too large',
|
||||
details: { reason: 'IMAGE_TOO_LARGE' },
|
||||
})),
|
||||
});
|
||||
|
||||
await assert.rejects(
|
||||
client.ask('session-fallback', imageContent(), { timeoutMs: 3_000 }),
|
||||
(error) => error.code === 'attachment-error' && error.details?.reason === 'IMAGE_TOO_LARGE',
|
||||
);
|
||||
assert.equal(promptCalls.length, 1);
|
||||
assert.equal(ingressUsed, false);
|
||||
});
|
||||
|
||||
test('a failed restage keeps the original model-rejection error for the user', async (t) => {
|
||||
const root = await workspace(t);
|
||||
const { client, promptCalls } = scriptedClient({
|
||||
workspaceRoot: root,
|
||||
fileIngressExecutor: async () => {
|
||||
throw new Error('ingress exploded');
|
||||
},
|
||||
onPrompt: (attempt) => (attempt === 1 ? Promise.reject(modelImageRejection()) : Promise.resolve({})),
|
||||
});
|
||||
|
||||
await assert.rejects(
|
||||
client.ask('session-fallback', imageContent(), { timeoutMs: 3_000 }),
|
||||
(error) => error.code === 'attachment-error'
|
||||
&& error.details?.reason === 'MODEL_DOES_NOT_SUPPORT_IMAGES',
|
||||
);
|
||||
assert.equal(promptCalls.length, 1);
|
||||
});
|
||||
|
||||
test('the fallback retry itself is never retried again', async (t) => {
|
||||
const root = await workspace(t);
|
||||
const { client, promptCalls } = scriptedClient({
|
||||
workspaceRoot: root,
|
||||
fileIngressExecutor: ({ files, workspace, signal }) => (
|
||||
stageInboundFiles({ files }, { workspace, signal })
|
||||
),
|
||||
onPrompt: () => Promise.reject(modelImageRejection()),
|
||||
});
|
||||
|
||||
await assert.rejects(
|
||||
client.ask('session-fallback', imageContent(), { timeoutMs: 3_000 }),
|
||||
(error) => error.code === 'attachment-error',
|
||||
);
|
||||
assert.equal(promptCalls.length, 2);
|
||||
});
|
||||
|
||||
test('text-only prompts keep the single-call behavior', async (t) => {
|
||||
const root = await workspace(t);
|
||||
const { client, promptCalls } = scriptedClient({
|
||||
workspaceRoot: root,
|
||||
onPrompt: () => Promise.resolve({}),
|
||||
});
|
||||
|
||||
assert.equal(
|
||||
await client.ask('session-fallback', 'plain question', { timeoutMs: 3_000 }),
|
||||
'answer from tools',
|
||||
);
|
||||
assert.equal(promptCalls.length, 1);
|
||||
assert.deepEqual(promptCalls[0].payload.content, [{ type: 'text', text: 'plain question' }]);
|
||||
});
|
||||
Loading…
Add table
Add a link
Reference in a new issue