mirror of
https://github.com/hansjone/oclaw.git
synced 2026-10-09 00:40:45 +08:00
feat(chat): image specialist session image fallback and WS UX
- Resolve legacy multimodal inputs from session history when the user sends text-only (assistant images first, then user uploads); optional env toggles. - Support relay_pointer image resolution in legacy lane collector. - Align uvicorn WebSocket frame limit with MAX_PAYLOAD_BYTES; document OCLAW_UVICORN_WS_MAX_SIZE and compatible-mode env knobs. - Admin chat: encode attachments after bubble preview; FileReader-based base64 for large files. - Tests and docs for new behavior. Co-authored-by: Cursor <cursoragent@cursor.com>
This commit is contained in:
parent
faeb067856
commit
396dba76ea
8 changed files with 368 additions and 28 deletions
|
|
@ -29,6 +29,11 @@
|
|||
AIA_ASSISTANT_GATEWAY_HOST=0.0.0.0
|
||||
AIA_ASSISTANT_GATEWAY_PORT=8787
|
||||
|
||||
# OCLAW_UVICORN_WS_MAX_SIZE uvicorn WebSocket 单帧最大字节;默认与 hello 里 maxPayload 一致(约 25MB)。
|
||||
# 不设时 python -m oclaw... 仍用旧 uvicorn 默认 16MB,多图 base64 一条 chat.send 易超限被断连,前端报 ws_closed。
|
||||
# 若用命令行启动 uvicorn,请自行加:--ws-max-size 26214400(或与 MAX_PAYLOAD_BYTES 一致)。
|
||||
# OCLAW_UVICORN_WS_MAX_SIZE=
|
||||
|
||||
# AIA_PREWARM_INTERVAL_SECONDS 后台预热任务周期(秒),默认 600,合法范围代码内 clamp。
|
||||
# 【前端】无。
|
||||
AIA_PREWARM_INTERVAL_SECONDS=600
|
||||
|
|
@ -287,6 +292,14 @@ AIA_IMAGE_EXPERT_MODEL=
|
|||
AIA_IMAGE_EXPERT_BASE_URL=
|
||||
AIA_IMAGE_EXPERT_API_KEY=
|
||||
AIA_IMAGE_EXPERT_CHAT_ENDPOINT=
|
||||
# compatible-mode defaults to OpenAI vision blocks (type + image_url/text). Set to 1 only if your gateway
|
||||
# requires DashScope doc shape without "type" on each content part (may break OpenAI-compatible gateways).
|
||||
# AIA_IMAGE_EXPERT_COMPAT_USE_DASHSCOPE_NATIVE_BLOCKS=
|
||||
# Max input images per turn (1–12; default 8).
|
||||
# AIA_IMAGE_EXPERT_MAX_INPUT_IMAGES=
|
||||
# Multi-turn image specialist: inherit images from assistant/user history when this turn has no attachment (0=off).
|
||||
# AIA_IMAGE_SPECIALIST_SESSION_IMAGE_FALLBACK=
|
||||
# AIA_IMAGE_SPECIALIST_FALLBACK_SCAN_MESSAGES=
|
||||
|
||||
# Optional; image expert top-level kwargs (parity with DashScope multimodal-gen samples)
|
||||
DASHSCOPE_IMAGE_STREAM=
|
||||
|
|
|
|||
|
|
@ -338,7 +338,27 @@
|
|||
|
||||
## 图片专家专线(``AIA_IMAGE_EXPERT_*``)
|
||||
|
||||
**路由 specialist=`image`** 时由 **`send_legacy_image_messages`** 调用;可走 native `{"image"}`/`{"text"}` 或 compatible-mode **`image_url` + `text` 块状**载荷。**不使用 `AIA_OCR_*`**,也不在图片专家链路继承主会话模型的 Base URL/API Key。
|
||||
**路由 specialist=`image`** 时由 **`send_legacy_image_messages`** 调用。**`compatible-mode/v1` + `/chat/completions`** 与 OpenAI 视觉接口一致:每条 `content[]` 必须含 **`type`**(默认使用 `image_url` + `text`)。若网关只吃 DashScope 无 `type` 的 `{"image"}`/`{"text}`,再设 **`AIA_IMAGE_EXPERT_COMPAT_USE_DASHSCOPE_NATIVE_BLOCKS=1`**。**不使用 `AIA_OCR_*`**。
|
||||
|
||||
- `AIA_IMAGE_EXPERT_COMPAT_USE_DASHSCOPE_NATIVE_BLOCKS`
|
||||
- 默认:未设置(等价关闭)
|
||||
- 作用:设为 `1` 时,在 **`compatible-mode`** Base URL 上仍发送 DashScope 文档形态(无 `type` 的 `image`/`text` 块)。仅当你的兼容网关明确支持该形态时使用;否则会遇到上游 `missing_required_parameter … content[n].type`。
|
||||
- 生效:`oclaw/platform/llm/image_legacy_client.py`
|
||||
|
||||
- `AIA_IMAGE_EXPERT_MAX_INPUT_IMAGES`
|
||||
- 默认:`8`(上限 `12`)
|
||||
- 作用:单轮发往模型的输入图数量上限(与 `collect_legacy_lane_images_from_attachments` / `send_legacy_image_messages` 裁剪一致)。
|
||||
- 生效:`oclaw/platform/llm/image_legacy_client.py`
|
||||
|
||||
- `AIA_IMAGE_SPECIALIST_SESSION_IMAGE_FALLBACK`
|
||||
- 默认:未设置(启用)
|
||||
- 作用:设为 `0` / `false` / `off` 时,**关闭** Chat 图片专家在多轮对话中「本轮无新图」时从会话历史继承图片(否则会在助手产出图与用户上传图之间回退)。关闭后与旧行为一致:无附件即提示不上图。
|
||||
- 生效:`oclaw/platform/llm/image_legacy_client.py`(经 `runtime/direct_loop.py` 的 `_maybe_image_specialist_legacy_gateway_turn` 调用)
|
||||
|
||||
- `AIA_IMAGE_SPECIALIST_FALLBACK_SCAN_MESSAGES`
|
||||
- 默认:`400`
|
||||
- 作用:历史回退时最多加载的最近消息条数(与 `SqliteStore.get_messages` 上限一致范围内)。
|
||||
- 生效:`oclaw/platform/llm/image_legacy_client.py`
|
||||
|
||||
- `AIA_IMAGE_EXPERT_BASE_URL`
|
||||
- 默认:无(必填,除非在代码中为 `send_legacy_image_messages(..., base_url=...)` 传入)
|
||||
|
|
@ -419,6 +439,12 @@
|
|||
- 作用:网关监听端口
|
||||
- 生效:`oclaw/app_server/fastapi_main.py`, `oclaw/runtime/operations/main.py`
|
||||
|
||||
- `OCLAW_UVICORN_WS_MAX_SIZE`
|
||||
- 默认:未设置时由 `oclaw.interfaces.http.fastapi_app:main` 使用与 `MAX_PAYLOAD_BYTES`(约 25MB)相同的值传入 uvicorn `ws_max_size`
|
||||
- 作用:**uvicorn 对 WebSocket 单帧的字节上限**(与 hello 里 `maxPayload` 应对齐)。uvicorn 自带默认约 **16MB**;两张高清图经 base64 塞进一条 `chat.send` 常超过 16MB,服务端直接断连,浏览器侧表现为 **`ws_closed`** / 请求失败
|
||||
- 说明:若用命令行 `uvicorn ...` 启动而未走 `fastapi_app.main()`,需自行加 `--ws-max-size` 或设置本变量(取决于你的启动入口是否读取)
|
||||
- 生效:`oclaw/interfaces/http/fastapi_app.py`
|
||||
|
||||
- `AIA_RUNTIME_LOG_DIR`
|
||||
- 默认:空(使用内部默认目录)
|
||||
- 作用:运行日志目录
|
||||
|
|
|
|||
|
|
@ -21,8 +21,8 @@
|
|||
## 2. 运行时数据流(网关 → 落库)
|
||||
|
||||
1. **`run_oclaw_direct_loop`** 在用户消息落库后立刻调用 **`_maybe_image_specialist_legacy_gateway_turn`**。
|
||||
2. **输入附件**:`collect_legacy_lane_images_from_attachments` 将 UI 附件规范为 `data:` URL 或 HTTP URL(`image_ref` / `input_image` / `image_url` 等)。
|
||||
3. **无图**:直接写入一条 `assistant` / `assistant_text` 提示语并返回,不调用上游。
|
||||
2. **输入附件**:`collect_legacy_lane_images_from_attachments` 将 UI 附件规范为 `data:` URL 或 HTTP URL(`image_ref` / `input_image` / `image_url` / `relay_pointer` 等)。若本轮无可用图且未关闭 **`AIA_IMAGE_SPECIALIST_SESSION_IMAGE_FALLBACK`**,则 **`collect_legacy_lane_images_with_session_fallback`** 按「最近一条带图的助手消息 → 更早的用户上传」从落库历史中补齐。**非 compatible** 的 native 网关:`messages[0].content` 为 DashScope 文档形态(若干 **`{"image": …}`** + **`{"text": …}`**)。**`compatible-mode/v1`** 默认改为 OpenAI 形状(每条含 **`type`**:`image_url` / `text`),否则上游常报缺少 `content[n].type`;仅当网关明确要求无 `type` 的旧形态时设 **`AIA_IMAGE_EXPERT_COMPAT_USE_DASHSCOPE_NATIVE_BLOCKS=1`**。
|
||||
3. **仍无图**(本轮与历史均未解析出输入图):直接写入一条 `assistant` / `assistant_text` 提示语并返回,不调用上游。
|
||||
4. **有图**:调用 **`send_legacy_image_messages`**(`/chat/completions` 兼容路径,非 Responses API)。
|
||||
5. **输出解析**:**`legacy_image_turn_bundle`**
|
||||
- 文本可为空;若有生成图则 **`materialize_legacy_response_output_attachments`** 写入本地 blob,产出 **`image_ref`**(或退化为 **`image_url`**)。
|
||||
|
|
|
|||
|
|
@ -17,6 +17,7 @@ const I18N = {
|
|||
"chat.noSessions": "暂无会话,点击新建开始。",
|
||||
"chat.loading": "加载中…",
|
||||
"chat.sending": "发送中…",
|
||||
"chat.encodingAttachments": "编码附件中…",
|
||||
"chat.error": "请求失败",
|
||||
"chat.loadMore": "加载更多",
|
||||
"chat.sessionMenu": "会话操作",
|
||||
|
|
@ -193,6 +194,7 @@ const I18N = {
|
|||
"chat.noSessions": "No sessions yet. Create one to start.",
|
||||
"chat.loading": "Loading…",
|
||||
"chat.sending": "Sending…",
|
||||
"chat.encodingAttachments": "Encoding attachments…",
|
||||
"chat.error": "Request failed",
|
||||
"chat.loadMore": "Load more",
|
||||
"chat.sessionMenu": "Session actions",
|
||||
|
|
@ -2588,12 +2590,17 @@ function syncAuthUserLabel() {
|
|||
}
|
||||
|
||||
async function fileToPayloadEntry(file) {
|
||||
const buf = await file.arrayBuffer();
|
||||
const bytes = new Uint8Array(buf);
|
||||
let binary = "";
|
||||
for (let i = 0; i < bytes.byteLength; i++) binary += String.fromCharCode(bytes[i]);
|
||||
const data_base64 = btoa(binary);
|
||||
return { name: file.name || "file", data_base64 };
|
||||
return new Promise((resolve, reject) => {
|
||||
const r = new FileReader();
|
||||
r.onload = () => {
|
||||
const dataUrl = String(r.result || "");
|
||||
const idx = dataUrl.indexOf(",");
|
||||
const data_base64 = idx >= 0 ? dataUrl.slice(idx + 1) : dataUrl;
|
||||
resolve({ name: file.name || "file", data_base64 });
|
||||
};
|
||||
r.onerror = () => reject(r.error || new Error("file_read_failed"));
|
||||
r.readAsDataURL(file);
|
||||
});
|
||||
}
|
||||
|
||||
async function renderLogin() {
|
||||
|
|
@ -4939,16 +4946,10 @@ ${autoLimit ? `<div style="margin-top:8px;"><span class="muted">auto-added claus
|
|||
btnSend.disabled = true;
|
||||
statusBar.textContent = t("chat.sending");
|
||||
let attPayload = null;
|
||||
if (filesSnapshot.length) {
|
||||
statusBar.textContent = currentLang === "zh" ? "准备附件中…" : "Preparing attachments…";
|
||||
attPayload = await Promise.all(filesSnapshot.map((f) => fileToPayloadEntry(f)));
|
||||
pendingFiles = [];
|
||||
syncPendingUi();
|
||||
}
|
||||
const userText =
|
||||
textRaw ||
|
||||
(hasFiles ? (currentLang === "zh" ? "(已上传附件)" : "(attachment uploaded)") : "");
|
||||
/** 发送过程中用本地 blob 预览,避免只显示文件名直到回合结束再拉历史 */
|
||||
/** 发送过程中用本地 blob 预览;先插入气泡再编码,避免大图 base64 阻塞主线程时长时间只有「准备附件」无预览 */
|
||||
const previewBlobUrls = [];
|
||||
try {
|
||||
textarea.value = "";
|
||||
|
|
@ -4974,6 +4975,13 @@ ${autoLimit ? `<div style="margin-top:8px;"><span class="muted">auto-added claus
|
|||
messagesEl.appendChild(userWrap);
|
||||
scrollMessagesToBottom(true);
|
||||
fitComposerTextarea();
|
||||
if (filesSnapshot.length) {
|
||||
await new Promise((resolve) => requestAnimationFrame(() => resolve()));
|
||||
statusBar.textContent = t("chat.encodingAttachments");
|
||||
attPayload = await Promise.all(filesSnapshot.map((f) => fileToPayloadEntry(f)));
|
||||
pendingFiles = [];
|
||||
syncPendingUi();
|
||||
}
|
||||
const turnId = `idem_${Date.now()}_${Math.random().toString(36).slice(2, 10)}`;
|
||||
statusBar.textContent = currentLang === "zh" ? "连接中…" : "Connecting…";
|
||||
const doneMeta = await sendMessageStream(userText, attPayload, turnId);
|
||||
|
|
|
|||
|
|
@ -20,6 +20,7 @@ from fastapi.staticfiles import StaticFiles
|
|||
from oclaw.runtime.application.gateway import process_inbound_payload_usecase
|
||||
from oclaw.interfaces.gateway.http_adapter import dispatch_gateway_http_method
|
||||
from oclaw.interfaces.ws import ws_gateway_loop
|
||||
from oclaw.interfaces.ws.common import MAX_PAYLOAD_BYTES
|
||||
from oclaw.interfaces.admin.routes import admin_static_dir, build_admin_router
|
||||
from oclaw.runtime.agents.agent_scope import list_agent_ids, resolve_agent_workspace_dir, resolve_default_agent_id
|
||||
from oclaw.interfaces.gateway.server_startup_plugins import prepare_gateway_plugin_bootstrap
|
||||
|
|
@ -305,7 +306,21 @@ def main() -> int:
|
|||
import uvicorn
|
||||
except Exception as exc:
|
||||
raise RuntimeError("missing dependency: uvicorn (pip install -r requirements.txt)") from exc
|
||||
uvicorn.run("oclaw.interfaces.http.fastapi_app:create_app", host=host, port=port, reload=False, factory=True)
|
||||
# Uvicorn default ws_max_size is 16MB; two photos as base64 in one chat.send often exceed that and the
|
||||
# server drops the socket (client sees Error: ws_closed). Default aligns with WS hello maxPayload.
|
||||
try:
|
||||
ws_max_size = int(os.getenv("OCLAW_UVICORN_WS_MAX_SIZE") or str(MAX_PAYLOAD_BYTES))
|
||||
except Exception:
|
||||
ws_max_size = int(MAX_PAYLOAD_BYTES)
|
||||
ws_max_size = max(1024, min(int(ws_max_size), 200_000_000))
|
||||
uvicorn.run(
|
||||
"oclaw.interfaces.http.fastapi_app:create_app",
|
||||
host=host,
|
||||
port=port,
|
||||
reload=False,
|
||||
factory=True,
|
||||
ws_max_size=ws_max_size,
|
||||
)
|
||||
return 0
|
||||
|
||||
|
||||
|
|
|
|||
|
|
@ -21,7 +21,7 @@ SDK-style extras:
|
|||
use ``DASHSCOPE_IMAGE_*`` env vars or JSON in ``AIA_IMAGE_EXPERT_REQUEST_EXTRA`` (alias: ``AIA_LEGACY_IMAGE_REQUEST_EXTRA``).
|
||||
|
||||
Compatibility roots:
|
||||
- For OpenAI-compat multimodal, use ``AIA_IMAGE_EXPERT_BASE_URL=https://dashscope.aliyuncs.com/compatible-mode/v1`` etc.
|
||||
- ``compatible-mode/v1`` **默认**使用 OpenAI Chat Completions 视觉块(每条含 ``type``:``image_url`` / ``text``);百炼兼容网关否则会报 ``missing_required_parameter … content[n].type``。若网关只吃 DashScope 文档形态(无 ``type`` 的 ``{"image"}``/``{"text"}``),设置 ``AIA_IMAGE_EXPERT_COMPAT_USE_DASHSCOPE_NATIVE_BLOCKS=1``。
|
||||
|
||||
For OpenAI-style ``image_url`` chat payloads (tool OCR / multimodal downgrade), use :mod:`oclaw.platform.llm.image_ocr_client`.
|
||||
|
||||
|
|
@ -33,6 +33,7 @@ from __future__ import annotations
|
|||
import base64
|
||||
import json
|
||||
import os
|
||||
import re
|
||||
import sys
|
||||
from typing import Any
|
||||
|
||||
|
|
@ -63,15 +64,46 @@ IMAGE_SPECIALIST_DEFAULT_PROMPT_ZH = (
|
|||
)
|
||||
|
||||
|
||||
def _max_legacy_input_images() -> int:
|
||||
try:
|
||||
v = int(os.getenv("AIA_IMAGE_EXPERT_MAX_INPUT_IMAGES") or "8")
|
||||
except Exception:
|
||||
v = 8
|
||||
return max(1, min(int(v), 12))
|
||||
|
||||
|
||||
def parse_message_attachments_json(raw: Any) -> list[dict[str, Any]]:
|
||||
"""Decode ``chat_message.attachments`` (JSON string, list, or dict) into attachment dict rows."""
|
||||
if raw is None:
|
||||
return []
|
||||
if isinstance(raw, list):
|
||||
return [x for x in raw if isinstance(x, dict)]
|
||||
if isinstance(raw, dict):
|
||||
return [raw]
|
||||
if isinstance(raw, str):
|
||||
s = raw.strip()
|
||||
if not s or s.lower() == "null":
|
||||
return []
|
||||
try:
|
||||
v = json.loads(s)
|
||||
if isinstance(v, list):
|
||||
return [x for x in v if isinstance(x, dict)]
|
||||
if isinstance(v, dict):
|
||||
return [v]
|
||||
except Exception:
|
||||
pass
|
||||
return []
|
||||
|
||||
|
||||
def collect_legacy_lane_images_from_attachments(
|
||||
attachments: list[dict[str, Any]] | None,
|
||||
*,
|
||||
max_images: int = 3,
|
||||
max_images: int | None = None,
|
||||
) -> list[str]:
|
||||
"""Normalize incoming UI/store attachments to URLs/data URLs for :func:`send_legacy_image_messages`."""
|
||||
from oclaw.platform.files.attachment_assets import attachment_id_to_data_url
|
||||
|
||||
cap = max(1, min(int(max_images), 12))
|
||||
cap = max(1, min(int(max_images if max_images is not None else _max_legacy_input_images()), 12))
|
||||
out: list[str] = []
|
||||
for att in attachments or []:
|
||||
if not isinstance(att, dict):
|
||||
|
|
@ -96,11 +128,91 @@ def collect_legacy_lane_images_from_attachments(
|
|||
u = str(att.get("url") or "").strip()
|
||||
if u:
|
||||
out.append(u)
|
||||
elif t == "relay_pointer":
|
||||
mime = str(att.get("mime") or "").strip().lower()
|
||||
if mime and not mime.startswith("image/"):
|
||||
continue
|
||||
uri = str(att.get("pointer_uri") or "").strip()
|
||||
aid = str(att.get("attachment_id") or "").strip()
|
||||
if not aid and uri:
|
||||
m = re.match(r"^relay://attachments/[^/]+/([a-f0-9]{8,128})$", uri, re.I)
|
||||
if m:
|
||||
aid = str(m.group(1)).lower()
|
||||
if aid:
|
||||
data_url = attachment_id_to_data_url(aid, mime=str(att.get("mime") or ""))
|
||||
if data_url:
|
||||
out.append(data_url)
|
||||
if len(out) >= cap:
|
||||
break
|
||||
return out
|
||||
|
||||
|
||||
def collect_legacy_lane_images_with_session_fallback(
|
||||
*,
|
||||
store: Any,
|
||||
session_id: str,
|
||||
attachments: list[dict[str, Any]] | None,
|
||||
max_images: int | None = None,
|
||||
) -> tuple[list[str], str]:
|
||||
"""Resolve legacy multimodal image URLs for the current turn, then session history.
|
||||
|
||||
When the current request has no embedded/URL images, scan persisted messages (newest→oldest):
|
||||
1) assistant rows with usable image attachments (e.g. last model output);
|
||||
2) user rows with images.
|
||||
|
||||
The trailing user bubble for this turn (text-only) is skipped when scanning.
|
||||
|
||||
Returns ``(images, source)`` where ``source`` is ``current``, ``assistant_history``,
|
||||
``user_history``, or empty when nothing was found.
|
||||
|
||||
Disable with ``AIA_IMAGE_SPECIALIST_SESSION_IMAGE_FALLBACK=0`` (default: on).
|
||||
"""
|
||||
cap = max_images if max_images is not None else _max_legacy_input_images()
|
||||
primary = collect_legacy_lane_images_from_attachments(attachments, max_images=cap)
|
||||
if primary:
|
||||
return primary, "current"
|
||||
if str(os.getenv("AIA_IMAGE_SPECIALIST_SESSION_IMAGE_FALLBACK") or "").strip().lower() in (
|
||||
"0",
|
||||
"false",
|
||||
"no",
|
||||
"off",
|
||||
):
|
||||
return [], ""
|
||||
|
||||
try:
|
||||
scan_n = int(os.getenv("AIA_IMAGE_SPECIALIST_FALLBACK_SCAN_MESSAGES") or "400")
|
||||
except Exception:
|
||||
scan_n = 400
|
||||
scan_n = max(1, min(scan_n, 2000))
|
||||
|
||||
try:
|
||||
msgs = store.get_messages(str(session_id or "").strip(), limit=scan_n)
|
||||
except Exception:
|
||||
msgs = []
|
||||
|
||||
if not msgs:
|
||||
return [], ""
|
||||
|
||||
hi = len(msgs) - 1
|
||||
if hi >= 0:
|
||||
last = msgs[hi]
|
||||
if str(getattr(last, "role", "") or "").strip().lower() == "user":
|
||||
la = parse_message_attachments_json(getattr(last, "attachments", None))
|
||||
if not collect_legacy_lane_images_from_attachments(la, max_images=1):
|
||||
hi -= 1
|
||||
|
||||
for phase in ("assistant", "user"):
|
||||
for i in range(hi, -1, -1):
|
||||
m = msgs[i]
|
||||
if str(getattr(m, "role", "") or "").strip().lower() != phase:
|
||||
continue
|
||||
la = parse_message_attachments_json(getattr(m, "attachments", None))
|
||||
got = collect_legacy_lane_images_from_attachments(la, max_images=cap)
|
||||
if got:
|
||||
return got, f"{phase}_history"
|
||||
return [], ""
|
||||
|
||||
|
||||
def normalize_legacy_output_image_urls(resp_images: Any, *, max_items: int = 12) -> list[str]:
|
||||
"""Flatten provider ``images`` / content parts to HTTP/data URLs (strings only)."""
|
||||
cap = max(1, min(int(max_items), 24))
|
||||
|
|
@ -340,7 +452,11 @@ def _dashscope_image_env_kw() -> dict[str, Any]:
|
|||
|
||||
|
||||
def _openai_compatible_vision_content(images: list[str], prompt_text: str) -> list[dict[str, Any]]:
|
||||
"""DashScope *compatible-mode* / OpenAI Chat Completions vision shape (NOT ``{"image":..., "text":...}``)."""
|
||||
"""OpenAI Chat Completions vision shape (``type`` + ``image_url`` / ``text``).
|
||||
|
||||
Required for typical ``compatible-mode/v1`` ``/chat/completions`` (each part must have ``type``).
|
||||
Image URLs may be ``https://…`` or ``data:image/…;base64,…``.
|
||||
"""
|
||||
blocks: list[dict[str, Any]] = []
|
||||
for img in images:
|
||||
blocks.append({"type": "image_url", "image_url": {"url": img}})
|
||||
|
|
@ -348,6 +464,16 @@ def _openai_compatible_vision_content(images: list[str], prompt_text: str) -> li
|
|||
return blocks
|
||||
|
||||
|
||||
def _compat_use_dashscope_native_blocks() -> bool:
|
||||
"""When ``1``, ``compatible-mode`` URLs use DashScope doc blocks without ``type`` (see :func:`_http_content_blocks`)."""
|
||||
return str(os.getenv("AIA_IMAGE_EXPERT_COMPAT_USE_DASHSCOPE_NATIVE_BLOCKS") or "").strip().lower() in (
|
||||
"1",
|
||||
"true",
|
||||
"yes",
|
||||
"on",
|
||||
)
|
||||
|
||||
|
||||
def _model_triggers_dashscope_native_fallback(model_name: str) -> bool:
|
||||
"""``qwen-image`` on OpenAI-compat ``/chat/completions`` often returns ``message.content=null``."""
|
||||
if str(os.getenv("AIA_IMAGE_EXPERT_DISABLE_DASHSCOPE_NATIVE_FALLBACK") or "").strip().lower() in (
|
||||
|
|
@ -408,7 +534,7 @@ def send_legacy_image_messages(
|
|||
if not images:
|
||||
return {"ok": False, "error": "at least one image input is required"}
|
||||
|
||||
raw_selected = [str(x).strip() for x in images if str(x).strip()][:3]
|
||||
raw_selected = [str(x).strip() for x in images if str(x).strip()][: _max_legacy_input_images()]
|
||||
selected: list[str] = []
|
||||
input_kind: list[str] = []
|
||||
for img in raw_selected:
|
||||
|
|
@ -423,10 +549,11 @@ def send_legacy_image_messages(
|
|||
if not selected:
|
||||
return {"ok": False, "error": "no usable image input (expected URL or data URL)"}
|
||||
|
||||
# Compatible-mode expects OpenAI-style ``image_url`` + ``text`` parts; DashScope-native HTTP uses plain ``{"image"}`` blocks.
|
||||
use_openai_blocks = "compatible-mode" in resolved_base_url.lower()
|
||||
# compatible-mode /chat/completions is OpenAI-shaped: every content part needs ``type`` (else 400 missing content[n].type).
|
||||
# Native DashScope HTTP (non-compatible base) uses ``{"image"}`` / ``{"text"}`` without ``type``.
|
||||
use_compatible_base = "compatible-mode" in resolved_base_url.lower()
|
||||
prompt_plain = str(prompt or "").strip() or render_prompt("image/default_edit_prompt.zh.md", strict=True)
|
||||
if use_openai_blocks:
|
||||
if use_compatible_base and not _compat_use_dashscope_native_blocks():
|
||||
content_multi = _openai_compatible_vision_content(selected, prompt_plain)
|
||||
else:
|
||||
content_multi = _http_content_blocks(selected, prompt, typed=False)
|
||||
|
|
@ -482,7 +609,7 @@ def send_legacy_image_messages(
|
|||
if (
|
||||
not str(text or "").strip()
|
||||
and not out_images
|
||||
and use_openai_blocks
|
||||
and use_compatible_base
|
||||
and _model_triggers_dashscope_native_fallback(model_name)
|
||||
):
|
||||
compat_extract_diag = build_extract_diag_empty(body)
|
||||
|
|
@ -596,6 +723,8 @@ def send_legacy_image_messages(
|
|||
__all__ = [
|
||||
"IMAGE_SPECIALIST_DEFAULT_PROMPT_ZH",
|
||||
"collect_legacy_lane_images_from_attachments",
|
||||
"collect_legacy_lane_images_with_session_fallback",
|
||||
"parse_message_attachments_json",
|
||||
"legacy_image_assistant_body_with_placeholder",
|
||||
"legacy_image_turn_bundle",
|
||||
"materialize_legacy_response_output_attachments",
|
||||
|
|
|
|||
|
|
@ -1134,13 +1134,17 @@ def _maybe_image_specialist_legacy_gateway_turn(
|
|||
|
||||
from oclaw.platform.llm.image_legacy_client import (
|
||||
IMAGE_SPECIALIST_DEFAULT_PROMPT_ZH,
|
||||
collect_legacy_lane_images_from_attachments,
|
||||
collect_legacy_lane_images_with_session_fallback,
|
||||
legacy_image_assistant_body_with_placeholder,
|
||||
legacy_image_turn_bundle,
|
||||
send_legacy_image_messages,
|
||||
)
|
||||
|
||||
imgs = collect_legacy_lane_images_from_attachments(attachments)
|
||||
imgs, legacy_img_src = collect_legacy_lane_images_with_session_fallback(
|
||||
store=store,
|
||||
session_id=session_id,
|
||||
attachments=attachments,
|
||||
)
|
||||
if not imgs:
|
||||
hint_en = "Image specialist received no image input. Attach an image and try again."
|
||||
hint_zh = "图片专家未收到可用的图片输入;请先上传或附上图片后再试。"
|
||||
|
|
@ -1160,6 +1164,11 @@ def _maybe_image_specialist_legacy_gateway_turn(
|
|||
)
|
||||
|
||||
if on_progress:
|
||||
if legacy_img_src.endswith("_history"):
|
||||
if str(lang or "").startswith("en"):
|
||||
on_progress("oclaw: reusing earlier session images (no new upload this turn)…")
|
||||
else:
|
||||
on_progress("oclaw: 本轮未上传新图,使用会话中较早的图片作为输入…")
|
||||
on_progress("oclaw: image specialist (legacy multimodal HTTP)…")
|
||||
|
||||
prompt_plain = str(user_text or "").strip()
|
||||
|
|
|
|||
|
|
@ -1,5 +1,9 @@
|
|||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
|
||||
import pytest
|
||||
|
||||
from oclaw.platform.llm.image_http_common import (
|
||||
dashscope_multimodal_http_ok,
|
||||
dashscope_native_multimodal_url_from_compatible_base,
|
||||
|
|
@ -8,10 +12,43 @@ from oclaw.platform.llm.image_http_common import (
|
|||
from oclaw.platform.llm.image_http_common import build_extract_diag_empty
|
||||
from oclaw.platform.llm.image_legacy_client import (
|
||||
collect_legacy_lane_images_from_attachments,
|
||||
collect_legacy_lane_images_with_session_fallback,
|
||||
legacy_image_assistant_body_with_placeholder,
|
||||
legacy_image_turn_bundle,
|
||||
normalize_legacy_output_image_urls,
|
||||
parse_message_attachments_json,
|
||||
)
|
||||
from oclaw.platform.llm.image_legacy_client import _http_content_blocks
|
||||
from oclaw.platform.llm.image_legacy_client import _openai_compatible_vision_content
|
||||
|
||||
|
||||
def test_openai_compatible_vision_has_type_per_part() -> None:
|
||||
"""compatible-mode /chat/completions expects type on each content element."""
|
||||
b = _openai_compatible_vision_content(
|
||||
["https://example.invalid/a.png", "data:image/png;base64,abcd"],
|
||||
"prompt",
|
||||
)
|
||||
assert len(b) == 3
|
||||
assert b[0]["type"] == "image_url" and "url" in b[0]["image_url"]
|
||||
assert b[1]["type"] == "image_url"
|
||||
assert b[2]["type"] == "text" and b[2]["text"] == "prompt"
|
||||
|
||||
|
||||
def test_http_content_blocks_multi_image_then_text() -> None:
|
||||
"""DashScope samples: one user message, content = N × {"image": url} then {"text": ...}."""
|
||||
b = _http_content_blocks(
|
||||
[
|
||||
"https://example.invalid/a.png",
|
||||
"https://example.invalid/b.png",
|
||||
],
|
||||
"合成说明",
|
||||
typed=False,
|
||||
)
|
||||
assert b == [
|
||||
{"image": "https://example.invalid/a.png"},
|
||||
{"image": "https://example.invalid/b.png"},
|
||||
{"text": "合成说明"},
|
||||
]
|
||||
|
||||
|
||||
def test_normalize_legacy_output_dict_image_parts() -> None:
|
||||
|
|
@ -246,3 +283,106 @@ def test_extract_text_and_images_openai_top_level_unchanged() -> None:
|
|||
)
|
||||
assert "hi" in text
|
||||
assert images == ["https://example.invalid/a.jpg"]
|
||||
|
||||
|
||||
def test_parse_message_attachments_json_string_list() -> None:
|
||||
raw = json.dumps([{"type": "image_url", "url": "https://example.invalid/z.png"}])
|
||||
got = parse_message_attachments_json(raw)
|
||||
assert len(got) == 1 and got[0]["url"] == "https://example.invalid/z.png"
|
||||
|
||||
|
||||
class _FakeHistMsg:
|
||||
__slots__ = ("role", "attachments")
|
||||
|
||||
def __init__(self, role: str, attachments: object) -> None:
|
||||
self.role = role
|
||||
self.attachments = attachments
|
||||
|
||||
|
||||
class _FakeHistStore:
|
||||
def __init__(self, msgs: list[_FakeHistMsg]) -> None:
|
||||
self._msgs = msgs
|
||||
|
||||
def get_messages(self, session_id: str, limit: int = 200) -> list[_FakeHistMsg]:
|
||||
_ = session_id
|
||||
return self._msgs[-limit:] if len(self._msgs) > limit else list(self._msgs)
|
||||
|
||||
|
||||
def test_session_fallback_prefers_latest_assistant_images() -> None:
|
||||
"""Newest assistant row with images wins over older user uploads."""
|
||||
msgs = [
|
||||
_FakeHistMsg(
|
||||
"user",
|
||||
json.dumps([{"type": "image_url", "url": "https://example.invalid/old-user.png"}]),
|
||||
),
|
||||
_FakeHistMsg(
|
||||
"assistant",
|
||||
json.dumps([{"type": "image_url", "url": "https://example.invalid/from-assistant.png"}]),
|
||||
),
|
||||
_FakeHistMsg("user", "null"),
|
||||
]
|
||||
store = _FakeHistStore(msgs)
|
||||
imgs, src = collect_legacy_lane_images_with_session_fallback(
|
||||
store=store,
|
||||
session_id="s1",
|
||||
attachments=[],
|
||||
)
|
||||
assert src == "assistant_history"
|
||||
assert imgs == ["https://example.invalid/from-assistant.png"]
|
||||
|
||||
|
||||
def test_session_fallback_user_history_when_no_assistant_images() -> None:
|
||||
msgs = [
|
||||
_FakeHistMsg(
|
||||
"user",
|
||||
json.dumps([{"type": "image_url", "url": "https://example.invalid/only-user.png"}]),
|
||||
),
|
||||
_FakeHistMsg("assistant", "[]"),
|
||||
_FakeHistMsg("user", "null"),
|
||||
]
|
||||
store = _FakeHistStore(msgs)
|
||||
imgs, src = collect_legacy_lane_images_with_session_fallback(
|
||||
store=store,
|
||||
session_id="s1",
|
||||
attachments=[],
|
||||
)
|
||||
assert src == "user_history"
|
||||
assert imgs == ["https://example.invalid/only-user.png"]
|
||||
|
||||
|
||||
def test_session_fallback_disabled_env(monkeypatch: pytest.MonkeyPatch) -> None:
|
||||
monkeypatch.setenv("AIA_IMAGE_SPECIALIST_SESSION_IMAGE_FALLBACK", "0")
|
||||
msgs = [
|
||||
_FakeHistMsg(
|
||||
"assistant",
|
||||
json.dumps([{"type": "image_url", "url": "https://example.invalid/a.png"}]),
|
||||
),
|
||||
_FakeHistMsg("user", "null"),
|
||||
]
|
||||
store = _FakeHistStore(msgs)
|
||||
imgs, src = collect_legacy_lane_images_with_session_fallback(
|
||||
store=store,
|
||||
session_id="s1",
|
||||
attachments=[],
|
||||
)
|
||||
assert src == ""
|
||||
assert imgs == []
|
||||
|
||||
|
||||
def test_session_fallback_current_attachments_skip_history(monkeypatch: pytest.MonkeyPatch) -> None:
|
||||
monkeypatch.delenv("AIA_IMAGE_SPECIALIST_SESSION_IMAGE_FALLBACK", raising=False)
|
||||
msgs = [
|
||||
_FakeHistMsg(
|
||||
"assistant",
|
||||
json.dumps([{"type": "image_url", "url": "https://example.invalid/hist.png"}]),
|
||||
),
|
||||
_FakeHistMsg("user", "null"),
|
||||
]
|
||||
store = _FakeHistStore(msgs)
|
||||
imgs, src = collect_legacy_lane_images_with_session_fallback(
|
||||
store=store,
|
||||
session_id="s1",
|
||||
attachments=[{"type": "image_url", "url": "https://example.invalid/current.png"}],
|
||||
)
|
||||
assert src == "current"
|
||||
assert imgs == ["https://example.invalid/current.png"]
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue