mirror of
https://github.com/hansjone/oclaw.git
synced 2026-10-08 23:33:16 +08:00
feat(video): DashScope video specialist with correct i2v request bodies
- Add video_generation_client: async video-synthesis, t2v vs i2v (Wan 2.7 input.media first_frame vs legacy img_url). - Coerce *-t2v* to *-i2v* when first frame present; AIA_VIDEO_I2V_INPUT_STYLE / I2V_MODEL overrides. - Gateway/direct_loop early return for video specialist; workspace video + factory allowlist. - Normalize WS and admin chat attachments (image_ref parity); mp4 attachment store roundtrip tests. Co-authored-by: Cursor <cursoragent@cursor.com>
This commit is contained in:
parent
398d85ba06
commit
4335040a27
21 changed files with 1152 additions and 56 deletions
|
|
@ -393,6 +393,67 @@
|
|||
|
||||
> **说明:** 历史上曾使用 `AIA_IMAGE_BASE_URL` / `AIA_IMAGE_API_KEY` / `AIA_IMAGE_MODEL` / `AIA_IMAGE_CHAT_ENDPOINT` 作为同一路由;当前实现 **不再读取** 上述变量作 OCR 通道配置,请统一改为 `AIA_OCR_*`。(`AIA_IMAGE_TOOL_RESULT_REPLAY_CAP_CHARS` 等为 **另一用途**,与 OCR 网关无关,仍保留原名。)
|
||||
|
||||
---
|
||||
|
||||
## 视频生成专家专线(``AIA_VIDEO_EXPERT_*``)
|
||||
|
||||
**路由 specialist=`video`** 时由 **`send_video_generation_request`** 调用百炼 / DashScope **异步文生视频** HTTP(``POST .../video-synthesis`` + ``GET .../api/v1/tasks/{task_id}``)。**区域**:北京 ``https://dashscope.aliyuncs.com``、新加坡 ``https://dashscope-intl.aliyuncs.com``、美东 ``https://dashscope-us.aliyuncs.com`` 等须与 API Key 一致。说明全文见 ``docs/VIDEO_SPECIALIST_LANE.md``。
|
||||
|
||||
- `AIA_VIDEO_SPECIALIST_DISABLE_LEGACY_GATEWAY_LANE`
|
||||
- 默认:未设置(关闭等价于 **启用** 专用 Early Return)
|
||||
- 作用:设为 `1` / `true` / `yes` / `on` 时,网关 **不再** 走专用视频 HTTP 线,改与普通 Chat 相同的模型环路。
|
||||
- 生效:`runtime/direct_loop.py`
|
||||
|
||||
- `AIA_VIDEO_EXPERT_BASE_URL`
|
||||
- 默认:无(必填,除非在调用处传入 `base_url`)
|
||||
- 作用:DashScope 根 URL(可为 ``compatible-mode/v1`` 前缀,实现会剥离后拼接原生路径)。
|
||||
- 生效:`oclaw/platform/llm/video_generation_client.py`
|
||||
|
||||
- `AIA_VIDEO_EXPERT_API_KEY`
|
||||
- 默认:无(必填,除非显式传 `api_key`)
|
||||
- 作用:``Authorization: Bearer`` API Key。
|
||||
- 生效:`oclaw/platform/llm/video_generation_client.py`
|
||||
|
||||
- `AIA_VIDEO_EXPERT_MODEL`
|
||||
- 默认:无(会话所选模型的 `model` 字段优先;无则读本变量)
|
||||
- 作用:Wan 模型 id:**文生视频**用 t2v 系列(如 ``wan2.2-t2v-plus``);**图生视频**(消息带首帧图时自动传 ``input.img_url``)须用 **i2v** 系列(如 ``wan2.6-i2v-flash``,以控制台为准)。
|
||||
- 生效:`oclaw/platform/llm/video_generation_client.py`
|
||||
|
||||
- `AIA_VIDEO_EXPERT_SYNTHESIS_PATH`
|
||||
- 默认:`api/v1/services/aigc/video-generation/video-synthesis`
|
||||
- 作用:相对 `AIA_VIDEO_EXPERT_BASE_URL` 剥离 compatible 后缀后的根路径拼接用。
|
||||
- 生效:`oclaw/platform/llm/video_generation_client.py`
|
||||
|
||||
- `AIA_VIDEO_EXPERT_POLL_INTERVAL_SEC`
|
||||
- 默认:`15`(限制在约 `3`~`120` 秒)
|
||||
- 作用:轮询任务状态间隔。
|
||||
- 生效:`oclaw/platform/llm/video_generation_client.py`
|
||||
|
||||
- `AIA_VIDEO_EXPERT_MAX_WAIT_SEC`
|
||||
- 默认:`900`(限制在约 `30`~`3600` 秒)
|
||||
- 作用:自提交任务起的最大等待时间;超时返回错误。
|
||||
- 生效:`oclaw/platform/llm/video_generation_client.py`
|
||||
|
||||
- `AIA_VIDEO_EXPERT_PARAMETERS_EXTRA`
|
||||
- 默认:无
|
||||
- 作用:JSON 对象,**浅合并**到请求体 `parameters`(后写入,故与 `DASHSCOPE_VIDEO_*` 同名键时 **以该 JSON 为准**)。
|
||||
- 生效:`oclaw/platform/llm/video_generation_client.py`
|
||||
|
||||
- `AIA_VIDEO_EXPERT_INPUT_EXTRA`
|
||||
- 默认:无
|
||||
- 作用:JSON 对象,合并到请求体 `input`(`prompt` 仍由运行时代码写入并覆盖同名键)。
|
||||
- 生效:`oclaw/platform/llm/video_generation_client.py`
|
||||
|
||||
- `AIA_VIDEO_EXPERT_DEBUG_PRINT_PAYLOAD`
|
||||
- 默认:未设置
|
||||
- 作用:设为 `1` 时在 stderr 打印提交 URL 与 JSON 请求体(勿在生产长期开启)。
|
||||
- 生效:`oclaw/platform/llm/video_generation_client.py`
|
||||
|
||||
- `DASHSCOPE_VIDEO_SIZE` / `DASHSCOPE_VIDEO_DURATION` / `DASHSCOPE_VIDEO_PROMPT_EXTEND` / `DASHSCOPE_VIDEO_NEGATIVE_PROMPT` / `DASHSCOPE_VIDEO_AUDIO_URL` / `DASHSCOPE_VIDEO_SHOT_TYPE` / `DASHSCOPE_VIDEO_WATERMARK` / `DASHSCOPE_VIDEO_SEED`
|
||||
- 默认:均未设置则不附加对应字段。
|
||||
- 作用:与官方示例字段对齐的便捷变量;细粒度控制可改用 `AIA_VIDEO_EXPERT_PARAMETERS_EXTRA` / `AIA_VIDEO_EXPERT_INPUT_EXTRA`。
|
||||
- 生效:`oclaw/platform/llm/video_generation_client.py`
|
||||
|
||||
## Memory / RAG
|
||||
|
||||
- `AIA_RAG_MODE`
|
||||
|
|
|
|||
|
|
@ -35,6 +35,11 @@ Transport selection happens in `oclaw/runtime/agents/factory.py`.
|
|||
- **Not** a separate profile transport: when the user selects specialist **`image`** in `/chat`, `runtime/direct_loop.py` takes an early return and calls **`platform/llm/image_legacy_client.send_legacy_image_messages`** (DashScope-style `/chat/completions`), so that turn does **not** use `OpenAIResponsesModel` / chat transports above.
|
||||
- Details, env vars, ACL, and UI hooks: **`docs/IMAGE_SPECIALIST_LANE.md`**.
|
||||
|
||||
### Chat UI「视频生成专家」(绕行本矩阵)
|
||||
|
||||
- When the user selects specialist **`video`**, `runtime/direct_loop.py` early-returns into **`platform/llm/video_generation_client.send_video_generation_request`** (DashScope async `video-synthesis` + task polling). Without an input image: **text-to-video**; with an image attachment (or session image fallback): **`input.img_url`** for **image-to-video** (use an i2v model id). That turn does **not** use the generic tool loop or `OpenAIResponsesModel`.
|
||||
- **`docs/VIDEO_SPECIALIST_LANE.md`**.
|
||||
|
||||
### `anthropic` (Anthropic Messages streaming)
|
||||
- **Transport**: `oclaw/platform/llm/transports/anthropic_messages.py::AnthropicMessagesModel`
|
||||
- **API**: Anthropic `messages.stream` surface (gateway must provide Anthropic-compatible protocol)
|
||||
|
|
|
|||
68
docs/VIDEO_SPECIALIST_LANE.md
Normal file
68
docs/VIDEO_SPECIALIST_LANE.md
Normal file
|
|
@ -0,0 +1,68 @@
|
|||
# 视频生成专家(Chat UI)专用链路
|
||||
|
||||
本文描述 **Admin `/chat` 选择「视频」专家**(工作区 id **`video`**)时的端到端路径。与图片专家类似,这是一条 **与通用 Responses / 主对话工具环路隔离** 的分支:在 `skill_binding_role == "video"` 时 **Early Return**,不进入常规 `run_oclaw_direct_loop` 多轮工具循环。
|
||||
|
||||
参考 API 形态:阿里云 Model Studio **Wan**(DashScope 异步:同一 `POST …/video-synthesis` → `GET …/api/v1/tasks/{task_id}`)。**文生视频**仅 `input.prompt`;**图生视频**额外设置 `input.img_url`(公网 HTTPS 或 `data:image/...;base64,...` 首帧),需使用 **i2v** 模型(如 `wan2.6-i2v-flash`,以控制台为准)。区域化的 **`base_url` 必须与 API Key 区域一致**。控制台 API 入口示例:[百炼控制台](https://bailian.console.aliyun.com/)。
|
||||
|
||||
---
|
||||
|
||||
## 1. 触发条件与隔离边界
|
||||
|
||||
| 条件 | 说明 |
|
||||
|------|------|
|
||||
| UI / Gateway | 用户选择专家 **`video`**,请求携带 `skill_binding_role` 为 **`video`**。 |
|
||||
| 入口守卫 | `runtime/direct_loop.py` 中 **`_maybe_video_specialist_legacy_gateway_turn`**:仅当 `skill_binding_role.lower() == "video"` 且未设置禁用开关时执行。 |
|
||||
| 禁用开关 | `AIA_VIDEO_SPECIALIST_DISABLE_LEGACY_GATEWAY_LANE=1`:关闭本 Early Return,视频专家改走与普通会话相同的模型/传输栈。 |
|
||||
|
||||
实现集中在 **`platform/llm/video_generation_client.py`**(HTTP + 轮询 + 附件落地),避免在 `openai_responses` 中分叉。
|
||||
|
||||
---
|
||||
|
||||
## 2. 运行时数据流(网关 → 落库)
|
||||
|
||||
1. **`run_oclaw_direct_loop`** 在用户消息落库后,在图片专家分支之后调用 **`_maybe_video_specialist_legacy_gateway_turn`**。
|
||||
2. **Prompt**:使用本轮用户文本;若为空则使用 **`VIDEO_SPECIALIST_DEFAULT_PROMPT_ZH`**(`video_generation_client`)。图生视频时 `prompt` 仍建议填写(描述期望动态与镜头)。
|
||||
3. **首帧图**:与图片专家相同,使用 **`collect_legacy_lane_images_with_session_fallback`**(`image_legacy_client`)从**本轮附件**或(未关闭 **`AIA_IMAGE_SPECIALIST_SESSION_IMAGE_FALLBACK`** 时)**会话历史**中取 **1 张**图,转为 URL / data URL 后写入 **`input.img_url`**。无图则走纯文生视频。
|
||||
4. **鉴权与根 URL**:优先使用会话所选模型的 **`model` / `base_url` / `api_key`**;缺省字段由 **`AIA_VIDEO_EXPERT_*`** 环境变量补全。若 `base_url` 指向 **`compatible-mode/v1`**,实现会剥离该后缀以拼接原生 DashScope 路径。
|
||||
5. **调用**:`send_video_generation_request` — `POST .../video-synthesis`(`X-DashScope-Async: enable`),再轮询 **`GET .../api/v1/tasks/{task_id}`** 直至 `SUCCEEDED` / 失败 / 超时。
|
||||
6. **输出**:成功时从 `output.video_url` 下载为本地 blob,产出 **`video_ref`**;下载失败时退化为仅带 **`url`** 的 `video_ref` 行(前端仍可尝试外链播放)。
|
||||
7. **占位文案**:`legacy_video_assistant_body_with_placeholder` 与图片专家对称(仅附件、无正文时插入中英文短句)。
|
||||
8. **编排**:`runtime/agents/specialist_agent.py` 在 `step.specialist == "video"` 时调用同一客户端(按父任务附件 + 父会话历史解析首帧),保证综合模式子专家与专家模式行为一致。
|
||||
|
||||
---
|
||||
|
||||
## 3. 执行器与工具面
|
||||
|
||||
- **`runtime/agents/factory.py`**:`video` 与 `image` 一样使用 **不可能工具名 allowlist**,避免默认工具注册混入。
|
||||
- **`runtime/gateway.py`**:综合模式 Manager 白名单包含 **`video`**,否则子专家选择会被回退。
|
||||
|
||||
---
|
||||
|
||||
## 4. 鉴权与附件(ACL)
|
||||
|
||||
与 **`image_ref`** 相同,助手消息上的 **`video_ref`** 依赖 **`attachment_acl`** 才能在严格模式下通过下载接口访问;落库时由 `SqliteStore.add_message` 链路处理。参见 **`docs/attachment-acl.md`**。
|
||||
|
||||
---
|
||||
|
||||
## 5. 前端
|
||||
|
||||
- **`interfaces/admin/static/chat.js`**:`specialistLabel` 对 **`video`** 显示短标签;`video_ref` 卡片在可用 blob URL 或外链 URL 时附加 **`<video controls>`** 便于预览。
|
||||
|
||||
---
|
||||
|
||||
## 6. 环境变量(索引)
|
||||
|
||||
详见 **`docs/ENVIRONMENT_VARIABLES.md`** 中 **`AIA_VIDEO_EXPERT_*`** 与 **`DASHSCOPE_VIDEO_*`** 小节。
|
||||
|
||||
---
|
||||
|
||||
## 7. 测试
|
||||
|
||||
- **`tests/test_video_generation_client.py`**:提交 / 轮询 HTTP 形态的单元测试(mock `httpx`)。
|
||||
|
||||
---
|
||||
|
||||
## 8. 变更原则
|
||||
|
||||
1. 默认只改 **`video_generation_client.py`**、`direct_loop` 的 **`_maybe_video_specialist_*`**、`specialist_agent` 视频分支、`factory` / `gateway` 白名单、**`chat.js`** 附件展示、本文与 **`ENVIRONMENT_VARIABLES.md`**。
|
||||
2. 勿在通用 **`openai_responses`** 中为视频专家单独绕路,除非产品明确要求统一传输。
|
||||
|
|
@ -2,7 +2,6 @@
|
|||
|
||||
from __future__ import annotations
|
||||
|
||||
import base64
|
||||
import json
|
||||
import queue
|
||||
import threading
|
||||
|
|
@ -28,8 +27,8 @@ from oclaw.platform.files.file_attachments import (
|
|||
DEFAULT_MAX_EXCEL_SHEETS,
|
||||
DEFAULT_TABULAR_ROWS_READ,
|
||||
clear_attachment_limits_cache,
|
||||
process_file_data,
|
||||
)
|
||||
from oclaw.interfaces.ws.common import normalize_ws_attachments
|
||||
from oclaw.platform.files.session_export import export_session_json, export_session_markdown
|
||||
from oclaw.platform.persistence.sqlite_store import SqliteStore
|
||||
from oclaw.runtime.gateway import OclawGateway
|
||||
|
|
@ -270,45 +269,9 @@ def _clear_stop(session_id: str) -> None:
|
|||
_CHAT_STOP_EVENTS.pop(str(session_id), None)
|
||||
|
||||
|
||||
def _decode_base64_payload(s: str) -> bytes | None:
|
||||
"""Decode base64 from browser (may lack padding; URL-safe variants)."""
|
||||
raw = (s or "").strip()
|
||||
if not raw:
|
||||
return None
|
||||
if raw.startswith("data:") and "," in raw:
|
||||
raw = raw.split(",", 1)[1].strip()
|
||||
raw = raw.replace("-", "+").replace("_", "/")
|
||||
pad = (-len(raw)) % 4
|
||||
if pad:
|
||||
raw += "=" * pad
|
||||
try:
|
||||
return base64.b64decode(raw, validate=False)
|
||||
except Exception:
|
||||
try:
|
||||
return base64.standard_b64decode(raw)
|
||||
except Exception:
|
||||
return None
|
||||
|
||||
|
||||
def _parse_attachments_payload(raw: Any) -> list[dict[str, Any]] | None:
|
||||
if not raw:
|
||||
return None
|
||||
if not isinstance(raw, list):
|
||||
return None
|
||||
out: list[dict[str, Any]] = []
|
||||
for item in raw:
|
||||
if not isinstance(item, dict):
|
||||
continue
|
||||
name = str(item.get("name") or "file").strip() or "file"
|
||||
b64 = item.get("data_base64") if "data_base64" in item else item.get("data")
|
||||
if not isinstance(b64, str) or not b64.strip():
|
||||
continue
|
||||
data = _decode_base64_payload(b64)
|
||||
if not data:
|
||||
continue
|
||||
got = process_file_data(name, data)
|
||||
if got:
|
||||
out.extend(got)
|
||||
"""Match WebSocket :func:`normalize_ws_attachments`: base64 uploads plus pass-through typed refs."""
|
||||
out = normalize_ws_attachments(raw)
|
||||
return out if out else None
|
||||
|
||||
|
||||
|
|
|
|||
|
|
@ -168,6 +168,7 @@ const I18N = {
|
|||
"chat.specialistGeneralistShort": "通用",
|
||||
"chat.specialistOpsShort": "运维",
|
||||
"chat.specialistImageShort": "图像",
|
||||
"chat.specialistVideoShort": "视频",
|
||||
"chat.specialistMemoryShort": "记忆",
|
||||
"chat.specialistManagerSelfShort": "全能者",
|
||||
"chat.attachment.download": "下载",
|
||||
|
|
@ -344,6 +345,7 @@ const I18N = {
|
|||
"chat.specialistGeneralistShort": "Generalist",
|
||||
"chat.specialistOpsShort": "Ops",
|
||||
"chat.specialistImageShort": "Image",
|
||||
"chat.specialistVideoShort": "Video",
|
||||
"chat.specialistMemoryShort": "Memory",
|
||||
"chat.specialistManagerSelfShort": "Manager",
|
||||
"chat.attachment.download": "Download",
|
||||
|
|
@ -649,6 +651,7 @@ function specialistLabel(specialist) {
|
|||
if (s === "manager_self") return t("chat.specialistManagerSelfShort");
|
||||
if (s === "ops") return t("chat.specialistOpsShort");
|
||||
if (s === "image") return t("chat.specialistImageShort");
|
||||
if (s === "video") return t("chat.specialistVideoShort");
|
||||
if (s === "memory") return t("chat.specialistMemoryShort");
|
||||
return t("chat.specialistGeneralistShort");
|
||||
}
|
||||
|
|
@ -2190,6 +2193,39 @@ async function renderAttachmentsEl(raw) {
|
|||
text: t("chat.attachment.download"),
|
||||
}),
|
||||
);
|
||||
if (String(typ || "") === "video_ref" && mime.toLowerCase().startsWith("video/")) {
|
||||
card.appendChild(
|
||||
el("video", {
|
||||
class: "chat-att-video",
|
||||
controls: true,
|
||||
preload: "metadata",
|
||||
style: "max-width:100%;max-height:420px;margin-top:8px;border-radius:8px;",
|
||||
src: url,
|
||||
}),
|
||||
);
|
||||
}
|
||||
}
|
||||
} else {
|
||||
const remote = String(att.url || "").trim();
|
||||
if (remote && String(typ || "") === "video_ref" && mime.toLowerCase().startsWith("video/")) {
|
||||
card.appendChild(
|
||||
el("video", {
|
||||
class: "chat-att-video",
|
||||
controls: true,
|
||||
preload: "metadata",
|
||||
style: "max-width:100%;max-height:420px;margin-top:8px;border-radius:8px;",
|
||||
src: remote,
|
||||
}),
|
||||
);
|
||||
card.appendChild(
|
||||
el("a", {
|
||||
class: "chat-att-ref__link",
|
||||
href: remote,
|
||||
target: "_blank",
|
||||
rel: "noopener noreferrer",
|
||||
text: t("chat.attachment.download"),
|
||||
}),
|
||||
);
|
||||
}
|
||||
}
|
||||
const canPreviewText = !!aid && (String(typ || "") === "text_ref" || isTextLikeMime(mime));
|
||||
|
|
|
|||
|
|
@ -111,7 +111,7 @@ async def run_agent_turn_via_bridge(
|
|||
|
||||
accepted_ms = now_ms()
|
||||
msg_text = str(p.get("message") or "").strip()
|
||||
attachments = list(p.get("attachments") or [])
|
||||
attachments = normalize_ws_attachments(p.get("attachments"))
|
||||
execution_mode = str(p.get("execution_mode") or "agent").strip().lower() or "agent"
|
||||
if execution_mode not in {"agent", "plan"}:
|
||||
execution_mode = "agent"
|
||||
|
|
|
|||
|
|
@ -8,11 +8,23 @@ import re
|
|||
import time
|
||||
from dataclasses import dataclass
|
||||
from pathlib import Path
|
||||
from typing import Any, Optional
|
||||
from typing import Any, Final, Optional
|
||||
|
||||
from oclaw.platform.config.paths import attachments_dir
|
||||
_META_SUFFIX: Final[str] = ".meta.json"
|
||||
_ATTACHMENT_ID_RE: Final[re.Pattern[str]] = re.compile(r"^[a-f0-9]{64}$")
|
||||
# Must match extensions produced by ``save_bytes`` so ``load_bytes`` / ``get_local_path`` / ``gc`` can find blobs.
|
||||
_ATTACHMENT_BLOB_EXTENSIONS: Final[tuple[str, ...]] = (
|
||||
".png",
|
||||
".jpg",
|
||||
".jpeg",
|
||||
".webp",
|
||||
".gif",
|
||||
".mp4",
|
||||
".webm",
|
||||
".mov",
|
||||
"", # extensionless fallback
|
||||
)
|
||||
|
||||
|
||||
def _utc_ts() -> int:
|
||||
|
|
@ -35,6 +47,12 @@ def _ext_from_mime(mime: str | None) -> str:
|
|||
return ".webp"
|
||||
if m == "image/gif":
|
||||
return ".gif"
|
||||
if m == "video/mp4":
|
||||
return ".mp4"
|
||||
if m == "video/webm":
|
||||
return ".webm"
|
||||
if m in ("video/quicktime", "video/mov"):
|
||||
return ".mov"
|
||||
return ""
|
||||
|
||||
|
||||
|
|
@ -48,6 +66,12 @@ def _guess_mime_from_ext(ext: str) -> str:
|
|||
return "image/webp"
|
||||
if e == "gif":
|
||||
return "image/gif"
|
||||
if e == "mp4":
|
||||
return "video/mp4"
|
||||
if e == "webm":
|
||||
return "video/webm"
|
||||
if e == "mov":
|
||||
return "video/quicktime"
|
||||
return "application/octet-stream"
|
||||
|
||||
|
||||
|
|
@ -132,7 +156,7 @@ class AttachmentAssetStore:
|
|||
blob = bytes(data or b"")
|
||||
h = hashlib.sha256(blob).hexdigest()
|
||||
ext = _ext_from_mime(mime) or (("." + filename.split(".")[-1].lower()) if "." in filename else "")
|
||||
ext = ext if ext in (".png", ".jpg", ".jpeg", ".webp", ".gif") else _ext_from_mime(mime) or ""
|
||||
ext = ext if ext in _ATTACHMENT_BLOB_EXTENSIONS[:-1] else _ext_from_mime(mime) or ""
|
||||
data_path = self._data_path(h, ext=ext)
|
||||
meta_path = self._meta_path(h)
|
||||
|
||||
|
|
@ -203,10 +227,9 @@ class AttachmentAssetStore:
|
|||
except Exception:
|
||||
return b"", None
|
||||
meta = self.get_meta(aid)
|
||||
# try find data file by scanning common extensions
|
||||
exts = (".png", ".jpg", ".jpeg", ".webp", ".gif", "")
|
||||
# try find data file by scanning known extensions
|
||||
data_path = None
|
||||
for ext in exts:
|
||||
for ext in _ATTACHMENT_BLOB_EXTENSIONS:
|
||||
p = self._data_path(aid, ext=ext)
|
||||
if p.exists():
|
||||
data_path = p
|
||||
|
|
@ -225,8 +248,7 @@ class AttachmentAssetStore:
|
|||
aid = self._normalize_attachment_id(attachment_id)
|
||||
except Exception:
|
||||
return None
|
||||
exts = (".png", ".jpg", ".jpeg", ".webp", ".gif", "")
|
||||
for ext in exts:
|
||||
for ext in _ATTACHMENT_BLOB_EXTENSIONS:
|
||||
p = self._data_path(aid, ext=ext)
|
||||
if p.exists():
|
||||
self.touch(aid)
|
||||
|
|
@ -287,7 +309,7 @@ class AttachmentAssetStore:
|
|||
mp.unlink(missing_ok=True) # py3.8+; on 3.14 ok
|
||||
except Exception:
|
||||
pass
|
||||
for ext in (".png", ".jpg", ".jpeg", ".webp", ".gif", ""):
|
||||
for ext in _ATTACHMENT_BLOB_EXTENSIONS:
|
||||
try:
|
||||
self._data_path(aid, ext=ext).unlink(missing_ok=True)
|
||||
except Exception:
|
||||
|
|
|
|||
466
platform/llm/video_generation_client.py
Normal file
466
platform/llm/video_generation_client.py
Normal file
|
|
@ -0,0 +1,466 @@
|
|||
"""DashScope / Model Studio **text-to-video** and **image-to-video** (async: create task → poll).
|
||||
|
||||
Used by the **video** specialist (Admin Chat + orchestration) to bypass the generic LLM tool loop,
|
||||
analogous to :mod:`oclaw.platform.llm.image_legacy_client` for the image specialist.
|
||||
|
||||
**Text-to-video:** ``input`` only ``prompt`` (plus optional ``audio_url`` / ``negative_prompt`` for some models).
|
||||
Use a **t2v** model id.
|
||||
|
||||
**Image-to-video (Wan 2.6 and earlier):** ``input.img_url`` — HTTPS or ``data:image/...;base64,...``.
|
||||
|
||||
**Image-to-video (Wan 2.7 general API):** ``input.media``, e.g.
|
||||
``[{"type": "first_frame", "url": "<https or data url>"}]``.
|
||||
Do **not** send ``img_url`` for that path (see Alibaba *image-to-video-general* API reference).
|
||||
|
||||
Override with ``AIA_VIDEO_I2V_INPUT_STYLE=media|img_url``; default is **auto** (``wan2.7`` in model id → ``media``, else ``img_url``).
|
||||
|
||||
HTTP contract (region of ``base_url`` must match the API key):
|
||||
|
||||
- ``POST {root}/api/v1/services/aigc/video-generation/video-synthesis`` with header
|
||||
``X-DashScope-Async: enable``
|
||||
- ``GET {root}/api/v1/tasks/{task_id}`` until ``output.task_status`` is terminal
|
||||
|
||||
References:
|
||||
- https://www.alibabacloud.com/help/en/model-studio/text-to-video-api-reference/
|
||||
- https://www.alibabacloud.com/help/en/model-studio/image-to-video-api-reference/
|
||||
- https://www.alibabacloud.com/help/en/model-studio/image-to-video-general-api-reference/
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import os
|
||||
import re
|
||||
import sys
|
||||
import time
|
||||
from typing import Any, Callable, Literal
|
||||
|
||||
import httpx
|
||||
|
||||
from oclaw.platform.files.attachment_assets import AttachmentAssetStore
|
||||
from oclaw.platform.llm.image_http_common import (
|
||||
compress_data_url_image,
|
||||
dashscope_multimodal_http_ok,
|
||||
download_http_url_bytes,
|
||||
is_data_url,
|
||||
join_url,
|
||||
post_with_retry,
|
||||
)
|
||||
|
||||
VIDEO_SPECIALIST_DEFAULT_PROMPT_ZH = "请根据描述生成一段短视频,画面稳定、主体清晰。"
|
||||
|
||||
_DEFAULT_SYNTHESIS_PATH = "api/v1/services/aigc/video-generation/video-synthesis"
|
||||
|
||||
|
||||
def _truthy_env(name: str) -> bool:
|
||||
return str(os.getenv(name) or "").strip().lower() in ("1", "true", "yes", "on")
|
||||
|
||||
|
||||
def env_video_expert_api_key() -> str:
|
||||
return (os.getenv("AIA_VIDEO_EXPERT_API_KEY") or "").strip()
|
||||
|
||||
|
||||
def env_video_expert_base_url() -> str:
|
||||
return (os.getenv("AIA_VIDEO_EXPERT_BASE_URL") or "").strip()
|
||||
|
||||
|
||||
def env_video_expert_model() -> str:
|
||||
return (os.getenv("AIA_VIDEO_EXPERT_MODEL") or "").strip()
|
||||
|
||||
|
||||
def _effective_video_model_for_request(model_name: str, *, img_url: str | None) -> str:
|
||||
"""DashScope Wan: **text-to-video** ids use ``*-t2v-*``; **image-to-video** needs ``*-i2v-*``.
|
||||
|
||||
Chat profiles often configure only a **t2v** model; if we still submit ``input.img_url``, the gateway
|
||||
may largely follow the prompt and ignore the photograph (looks like unrelated output). When a first
|
||||
frame is present we therefore coerce ``-t2v-`` → ``-i2v-`` unless disabled or overridden.
|
||||
"""
|
||||
m = (model_name or "").strip()
|
||||
if not m:
|
||||
return m
|
||||
if not str(img_url or "").strip():
|
||||
return m
|
||||
if _truthy_env("AIA_VIDEO_EXPERT_DISABLE_I2V_MODEL_COERCION"):
|
||||
return m
|
||||
override = (os.getenv("AIA_VIDEO_EXPERT_I2V_MODEL") or "").strip()
|
||||
if override:
|
||||
return override
|
||||
if "-t2v-" in m:
|
||||
coerced = m.replace("-t2v-", "-i2v-", 1)
|
||||
if coerced != m:
|
||||
return coerced
|
||||
low = m.lower().rstrip()
|
||||
if low.endswith("-t2v"):
|
||||
coerced = m[: -len("-t2v")] + "-i2v"
|
||||
if coerced != m:
|
||||
return coerced
|
||||
return m
|
||||
|
||||
|
||||
def _i2v_first_frame_input_style(*, model_name: str) -> Literal["img_url", "media"]:
|
||||
"""How to pass the first-frame image: legacy ``img_url`` vs Wan 2.7 ``media`` array."""
|
||||
raw = (os.getenv("AIA_VIDEO_I2V_INPUT_STYLE") or "").strip().lower()
|
||||
if raw in ("media", "wan27", "general"):
|
||||
return "media"
|
||||
if raw in ("img_url", "legacy", "wan26"):
|
||||
return "img_url"
|
||||
ml = (model_name or "").strip().lower()
|
||||
return "media" if "wan2.7" in ml else "img_url"
|
||||
|
||||
|
||||
def env_video_expert_synthesis_path() -> str:
|
||||
raw = (os.getenv("AIA_VIDEO_EXPERT_SYNTHESIS_PATH") or "").strip().lstrip("/")
|
||||
return raw or _DEFAULT_SYNTHESIS_PATH
|
||||
|
||||
|
||||
def _poll_interval_sec() -> float:
|
||||
raw = (os.getenv("AIA_VIDEO_EXPERT_POLL_INTERVAL_SEC") or "").strip()
|
||||
if not raw:
|
||||
return 15.0
|
||||
try:
|
||||
return max(3.0, min(float(raw), 120.0))
|
||||
except ValueError:
|
||||
return 15.0
|
||||
|
||||
|
||||
def _max_wait_sec() -> float:
|
||||
raw = (os.getenv("AIA_VIDEO_EXPERT_MAX_WAIT_SEC") or "").strip()
|
||||
if not raw:
|
||||
return 900.0
|
||||
try:
|
||||
return max(30.0, min(float(raw), 3600.0))
|
||||
except ValueError:
|
||||
return 900.0
|
||||
|
||||
|
||||
def dashscope_api_root_from_base_url(base_url: str) -> str:
|
||||
"""Strip OpenAI-compatible suffixes so native DashScope paths can be joined."""
|
||||
b = (base_url or "").strip().rstrip("/")
|
||||
if not b:
|
||||
return ""
|
||||
root = re.sub(r"/compatible-mode/v\d+(?:/.*)?$", "", b, flags=re.IGNORECASE).rstrip("/")
|
||||
if not root or root == b:
|
||||
root = re.sub(r"/compatible-mode/?$", "", b, flags=re.IGNORECASE).rstrip("/")
|
||||
return root or b
|
||||
|
||||
|
||||
def _parameters_extra_from_env() -> dict[str, Any]:
|
||||
raw = (os.getenv("AIA_VIDEO_EXPERT_PARAMETERS_EXTRA") or "").strip()
|
||||
if not raw:
|
||||
return {}
|
||||
try:
|
||||
v = json.loads(raw)
|
||||
return v if isinstance(v, dict) else {}
|
||||
except Exception:
|
||||
return {}
|
||||
|
||||
|
||||
def _input_extra_from_env() -> dict[str, Any]:
|
||||
raw = (os.getenv("AIA_VIDEO_EXPERT_INPUT_EXTRA") or "").strip()
|
||||
if not raw:
|
||||
return {}
|
||||
try:
|
||||
v = json.loads(raw)
|
||||
return v if isinstance(v, dict) else {}
|
||||
except Exception:
|
||||
return {}
|
||||
|
||||
|
||||
def _dashscope_video_env_parameters() -> dict[str, Any]:
|
||||
out: dict[str, Any] = {}
|
||||
size = (os.getenv("DASHSCOPE_VIDEO_SIZE") or "").strip()
|
||||
if size:
|
||||
out["size"] = size
|
||||
raw_dur = (os.getenv("DASHSCOPE_VIDEO_DURATION") or "").strip()
|
||||
if raw_dur.isdigit():
|
||||
out["duration"] = max(1, min(int(raw_dur), 600))
|
||||
pe = (os.getenv("DASHSCOPE_VIDEO_PROMPT_EXTEND") or "").strip().lower()
|
||||
if pe in ("1", "true", "yes", "on"):
|
||||
out["prompt_extend"] = True
|
||||
elif pe in ("0", "false", "no", "off"):
|
||||
out["prompt_extend"] = False
|
||||
st = (os.getenv("DASHSCOPE_VIDEO_SHOT_TYPE") or "").strip()
|
||||
if st:
|
||||
out["shot_type"] = st
|
||||
wm = (os.getenv("DASHSCOPE_VIDEO_WATERMARK") or "").strip().lower()
|
||||
if wm in ("1", "true", "yes", "on"):
|
||||
out["watermark"] = True
|
||||
elif wm in ("0", "false", "no", "off"):
|
||||
out["watermark"] = False
|
||||
seed = (os.getenv("DASHSCOPE_VIDEO_SEED") or "").strip()
|
||||
if seed.isdigit():
|
||||
out["seed"] = max(0, min(int(seed), 2_147_483_647))
|
||||
return out
|
||||
|
||||
|
||||
def _dashscope_video_env_input() -> dict[str, Any]:
|
||||
out: dict[str, Any] = {}
|
||||
neg = os.getenv("DASHSCOPE_VIDEO_NEGATIVE_PROMPT")
|
||||
if neg is not None and str(neg).strip():
|
||||
out["negative_prompt"] = str(neg)
|
||||
audio = (os.getenv("DASHSCOPE_VIDEO_AUDIO_URL") or "").strip()
|
||||
if audio:
|
||||
out["audio_url"] = audio
|
||||
return out
|
||||
|
||||
|
||||
def _extract_output(body: dict[str, Any]) -> dict[str, Any]:
|
||||
out = body.get("output")
|
||||
return out if isinstance(out, dict) else {}
|
||||
|
||||
|
||||
def _stderr_debug_video(url: str, payload: dict[str, Any]) -> None:
|
||||
if not _truthy_env("AIA_VIDEO_EXPERT_DEBUG_PRINT_PAYLOAD"):
|
||||
return
|
||||
try:
|
||||
txt = json.dumps(payload, ensure_ascii=False, indent=2, default=str)
|
||||
sys.stderr.write(f"\n[oclaw video_generation] POST {url}\n{txt}\n\n")
|
||||
sys.stderr.flush()
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
|
||||
def send_video_generation_request(
|
||||
*,
|
||||
prompt: str,
|
||||
model: str | None = None,
|
||||
api_key: str | None = None,
|
||||
base_url: str | None = None,
|
||||
img_url: str | None = None,
|
||||
timeout_submit_sec: float = 120.0,
|
||||
on_progress: Callable[[str], None] | None = None,
|
||||
should_stop: Callable[[], bool] | None = None,
|
||||
) -> dict[str, Any]:
|
||||
"""Create an async video task and poll until terminal state or timeout.
|
||||
|
||||
When ``img_url`` is set (HTTPS URL or data URL), it becomes the first-frame reference: for Wan **2.7**
|
||||
i2v models this is sent as ``input.media`` with ``type: first_frame`` and ``url``; for older i2v models
|
||||
as ``input.img_url``. Caller image overrides any overlapping keys from ``AIA_VIDEO_EXPERT_INPUT_EXTRA`` after merge
|
||||
(style selected by model id and ``AIA_VIDEO_I2V_INPUT_STYLE``).
|
||||
"""
|
||||
resolved_key = (api_key or env_video_expert_api_key()).strip()
|
||||
raw_base = (base_url or env_video_expert_base_url()).strip()
|
||||
root = dashscope_api_root_from_base_url(raw_base)
|
||||
model_name = ((model or "").strip() or env_video_expert_model()).strip()
|
||||
if not resolved_key or not root:
|
||||
return {
|
||||
"ok": False,
|
||||
"error": "missing AIA_VIDEO_EXPERT_API_KEY or AIA_VIDEO_EXPERT_BASE_URL (or pass api_key and base_url)",
|
||||
}
|
||||
if not model_name:
|
||||
return {
|
||||
"ok": False,
|
||||
"error": "missing model (chosen profile/model=… or set AIA_VIDEO_EXPERT_MODEL)",
|
||||
}
|
||||
prompt_plain = str(prompt or "").strip()
|
||||
if not prompt_plain:
|
||||
prompt_plain = VIDEO_SPECIALIST_DEFAULT_PROMPT_ZH
|
||||
|
||||
model_name = _effective_video_model_for_request(model_name, img_url=img_url)
|
||||
|
||||
path = env_video_expert_synthesis_path()
|
||||
url_submit = join_url(root, path)
|
||||
input_body: dict[str, Any] = {
|
||||
**_input_extra_from_env(),
|
||||
**_dashscope_video_env_input(),
|
||||
"prompt": prompt_plain,
|
||||
}
|
||||
iu = str(img_url or "").strip()
|
||||
if iu:
|
||||
if is_data_url(iu):
|
||||
iu = compress_data_url_image(iu)
|
||||
_style = _i2v_first_frame_input_style(model_name=model_name)
|
||||
if _style == "media":
|
||||
input_body.pop("img_url", None)
|
||||
first = {"type": "first_frame", "url": iu}
|
||||
prev = input_body.get("media")
|
||||
rest: list[dict[str, Any]] = []
|
||||
if isinstance(prev, list):
|
||||
for item in prev:
|
||||
if isinstance(item, dict) and str(item.get("type") or "").strip().lower() != "first_frame":
|
||||
rest.append(dict(item))
|
||||
input_body["media"] = [first, *rest]
|
||||
else:
|
||||
input_body.pop("media", None)
|
||||
input_body["img_url"] = iu
|
||||
parameters: dict[str, Any] = {
|
||||
**_dashscope_video_env_parameters(),
|
||||
**_parameters_extra_from_env(),
|
||||
}
|
||||
payload = {"model": model_name, "input": input_body, "parameters": parameters}
|
||||
headers = {
|
||||
"Authorization": f"Bearer {resolved_key}",
|
||||
"Content-Type": "application/json",
|
||||
"X-DashScope-Async": "enable",
|
||||
}
|
||||
_stderr_debug_video(url_submit, payload)
|
||||
|
||||
submit_timeout = httpx.Timeout(max(30.0, float(timeout_submit_sec)), connect=45.0)
|
||||
with httpx.Client(timeout=submit_timeout, follow_redirects=True) as client:
|
||||
try:
|
||||
r = post_with_retry(client, url=url_submit, headers=headers, payload=payload)
|
||||
except Exception as e:
|
||||
return {"ok": False, "error": f"submit failed: {type(e).__name__}: {e}"}
|
||||
if r.status_code >= 400:
|
||||
return {"ok": False, "error": f"submit http {r.status_code}: {r.text[:800]}"}
|
||||
try:
|
||||
body = r.json()
|
||||
except Exception:
|
||||
return {"ok": False, "error": f"submit non-json: {r.text[:500]}"}
|
||||
if not isinstance(body, dict):
|
||||
return {"ok": False, "error": "submit response is not a JSON object"}
|
||||
ok0, err0 = dashscope_multimodal_http_ok(body)
|
||||
if not ok0:
|
||||
return {"ok": False, "error": err0 or "submit rejected by provider"}
|
||||
out0 = _extract_output(body)
|
||||
task_id = str(out0.get("task_id") or "").strip()
|
||||
st0 = str(out0.get("task_status") or "").strip().upper()
|
||||
v0 = str(out0.get("video_url") or "").strip()
|
||||
if st0 == "SUCCEEDED" and v0.startswith(("http://", "https://")):
|
||||
return {"ok": True, "video_urls": [v0], "task_id": task_id or None}
|
||||
if st0 in {"FAILED", "UNKNOWN"}:
|
||||
msg = str(out0.get("message") or out0.get("msg") or "").strip()
|
||||
code = str(out0.get("code") or "").strip()
|
||||
bit = f"{code}: {msg}".strip(": ").strip() or st0
|
||||
return {"ok": False, "error": bit, "task_id": task_id or None}
|
||||
if not task_id:
|
||||
return {"ok": False, "error": "submit response missing output.task_id", "raw": body}
|
||||
|
||||
poll_interval = _poll_interval_sec()
|
||||
max_wait = _max_wait_sec()
|
||||
deadline = time.monotonic() + max_wait
|
||||
task_url = join_url(root, f"api/v1/tasks/{task_id}")
|
||||
headers_get = {"Authorization": f"Bearer {resolved_key}"}
|
||||
poll_timeout = httpx.Timeout(120.0, connect=45.0)
|
||||
with httpx.Client(timeout=poll_timeout, follow_redirects=True) as client:
|
||||
while time.monotonic() < deadline:
|
||||
if should_stop and should_stop():
|
||||
return {"ok": False, "error": "stopped", "task_id": task_id}
|
||||
try:
|
||||
gr = client.get(task_url, headers=headers_get)
|
||||
except Exception as e:
|
||||
return {"ok": False, "error": f"poll failed: {type(e).__name__}: {e}", "task_id": task_id}
|
||||
if gr.status_code >= 400:
|
||||
return {
|
||||
"ok": False,
|
||||
"error": f"poll http {gr.status_code}: {gr.text[:800]}",
|
||||
"task_id": task_id,
|
||||
}
|
||||
try:
|
||||
tb = gr.json()
|
||||
except Exception:
|
||||
return {"ok": False, "error": f"poll non-json: {gr.text[:500]}", "task_id": task_id}
|
||||
if not isinstance(tb, dict):
|
||||
return {"ok": False, "error": "poll response not a JSON object", "task_id": task_id}
|
||||
okp, errp = dashscope_multimodal_http_ok(tb)
|
||||
if not okp:
|
||||
return {"ok": False, "error": errp or "poll rejected", "task_id": task_id}
|
||||
tout = _extract_output(tb)
|
||||
st = str(tout.get("task_status") or "").strip().upper()
|
||||
vurl = str(tout.get("video_url") or "").strip()
|
||||
if st == "SUCCEEDED" and vurl.startswith(("http://", "https://")):
|
||||
return {"ok": True, "video_urls": [vurl], "task_id": task_id}
|
||||
if st in {"FAILED", "UNKNOWN", "CANCELED"}:
|
||||
msg = str(tout.get("message") or tout.get("msg") or "").strip()
|
||||
code = str(tout.get("code") or "").strip()
|
||||
bit = f"{code}: {msg}".strip(": ").strip() or st
|
||||
return {"ok": False, "error": bit, "task_id": task_id}
|
||||
if on_progress and st in {"PENDING", "RUNNING"}:
|
||||
on_progress(f"oclaw: video task {st.lower()}…")
|
||||
time.sleep(poll_interval)
|
||||
|
||||
return {"ok": False, "error": f"poll timeout after {int(max_wait)}s", "task_id": task_id}
|
||||
|
||||
|
||||
def materialize_video_output_attachments(
|
||||
video_urls: Any,
|
||||
*,
|
||||
max_videos: int = 1,
|
||||
) -> list[dict[str, Any]]:
|
||||
"""Download remote MP4 (etc.) into ``video_ref`` rows; on failure fall back to best-effort metadata."""
|
||||
cap = max(1, min(int(max_videos), 4))
|
||||
produced: list[dict[str, Any]] = []
|
||||
if not isinstance(video_urls, list):
|
||||
return produced
|
||||
urls = [str(u).strip() for u in video_urls if str(u).strip().startswith(("http://", "https://"))]
|
||||
|
||||
store = AttachmentAssetStore()
|
||||
ua = "Mozilla/5.0 (compatible; oclaw-video-expert/1.0; +https://github.com/)"
|
||||
for idx, u in enumerate(urls[:cap], start=1):
|
||||
try:
|
||||
blob, ctype = download_http_url_bytes(u, user_agent=ua)
|
||||
if not blob:
|
||||
raise ValueError("empty body")
|
||||
mime = (ctype.split(";", 1)[0].strip() if ctype else "") or "video/mp4"
|
||||
ext = ".mp4"
|
||||
if mime == "video/webm":
|
||||
ext = ".webm"
|
||||
elif mime in ("video/quicktime", "video/mov"):
|
||||
ext = ".mov"
|
||||
meta = store.save_bytes(blob, filename=f"video-output-{idx}{ext}", mime=mime)
|
||||
produced.append(
|
||||
{
|
||||
"type": "video_ref",
|
||||
"attachment_id": meta.attachment_id,
|
||||
"name": meta.name,
|
||||
"mime": meta.mime,
|
||||
"bytes": meta.bytes,
|
||||
}
|
||||
)
|
||||
except Exception:
|
||||
produced.append(
|
||||
{
|
||||
"type": "video_ref",
|
||||
"name": f"video-output-{idx}.mp4",
|
||||
"mime": "video/mp4",
|
||||
"url": u,
|
||||
}
|
||||
)
|
||||
return produced
|
||||
|
||||
|
||||
def legacy_video_turn_bundle(resp: dict[str, Any]) -> tuple[bool, str, list[dict[str, Any]]]:
|
||||
ok = bool(resp.get("ok"))
|
||||
text = str(resp.get("text") or "").strip()
|
||||
if not ok:
|
||||
err = str(resp.get("error") or "").strip()
|
||||
return False, f"Video generation failed: {err or 'unknown error'}", []
|
||||
urls = resp.get("video_urls")
|
||||
if not isinstance(urls, list):
|
||||
urls = []
|
||||
atts = materialize_video_output_attachments(urls, max_videos=1)
|
||||
if atts:
|
||||
return True, text, atts
|
||||
if text:
|
||||
return True, text, []
|
||||
return False, "Video specialist failed: empty response from provider.", []
|
||||
|
||||
|
||||
def legacy_video_assistant_body_with_placeholder(
|
||||
*,
|
||||
lang: str | None,
|
||||
body_text: str,
|
||||
produced: list[dict[str, Any]] | None,
|
||||
) -> str:
|
||||
if str(body_text or "").strip():
|
||||
return str(body_text or "")
|
||||
if produced:
|
||||
return (
|
||||
"Generated video (see attachment below)."
|
||||
if str(lang or "").startswith("en")
|
||||
else "已生成视频(见下方附件)。"
|
||||
)
|
||||
return str(body_text or "")
|
||||
|
||||
|
||||
__all__ = [
|
||||
"VIDEO_SPECIALIST_DEFAULT_PROMPT_ZH",
|
||||
"dashscope_api_root_from_base_url",
|
||||
"env_video_expert_api_key",
|
||||
"env_video_expert_base_url",
|
||||
"env_video_expert_model",
|
||||
"legacy_video_assistant_body_with_placeholder",
|
||||
"legacy_video_turn_bundle",
|
||||
"materialize_video_output_attachments",
|
||||
"send_video_generation_request",
|
||||
]
|
||||
|
|
@ -276,7 +276,7 @@ def build_ops_agent(
|
|||
|
||||
|
||||
# `default_registry` treats empty allow_tags + empty allow_tools as "no filter". Use an impossible
|
||||
# tool name so the image specialist gets an empty tool surface (vision-only turns).
|
||||
# tool name so image/video specialists get an empty tool surface (dedicated HTTP lanes).
|
||||
_IMAGE_SPECIALIST_TOOL_ALLOWLIST: tuple[str, ...] = ("__oclaw_image_specialist_no_tools__",)
|
||||
|
||||
|
||||
|
|
@ -334,7 +334,7 @@ def build_gateway_executor(
|
|||
"path_policy_user_id": path_policy_user_id,
|
||||
"store": store,
|
||||
}
|
||||
if prof.name == "image":
|
||||
if prof.name in {"image", "video"}:
|
||||
reg_kw["allow_tools"] = list(_IMAGE_SPECIALIST_TOOL_ALLOWLIST)
|
||||
tools = default_registry(**reg_kw)
|
||||
return Agent(
|
||||
|
|
|
|||
|
|
@ -16,10 +16,17 @@ from oclaw.platform.persistence.sqlite_store import SqliteStore
|
|||
from oclaw.platform.llm.image_legacy_client import (
|
||||
IMAGE_SPECIALIST_DEFAULT_PROMPT_ZH,
|
||||
collect_legacy_lane_images_from_attachments,
|
||||
collect_legacy_lane_images_with_session_fallback,
|
||||
legacy_image_assistant_body_with_placeholder,
|
||||
legacy_image_turn_bundle,
|
||||
send_legacy_image_messages,
|
||||
)
|
||||
from oclaw.platform.llm.video_generation_client import (
|
||||
VIDEO_SPECIALIST_DEFAULT_PROMPT_ZH,
|
||||
legacy_video_assistant_body_with_placeholder,
|
||||
legacy_video_turn_bundle,
|
||||
send_video_generation_request,
|
||||
)
|
||||
from oclaw.runtime.tools import default_registry
|
||||
from oclaw.runtime.agents.specialists import expert_name_for_specialist
|
||||
|
||||
|
|
@ -301,6 +308,51 @@ class SpecialistAgentRunner:
|
|||
tool_traces=(),
|
||||
notes="image_pipeline",
|
||||
)
|
||||
elif step.specialist == "video":
|
||||
image_protocol = "video_generation.http"
|
||||
_, chosen_model, _ = self._resolve_profile_and_model(step.specialist)
|
||||
text_parts = [
|
||||
str(x).strip()
|
||||
for x in (step.objective, step.input_text, parent_task.user_text)
|
||||
if str(x or "").strip()
|
||||
]
|
||||
user_text_v = "\n".join(text_parts) if text_parts else VIDEO_SPECIALIST_DEFAULT_PROMPT_ZH
|
||||
policy_sid = str(policy_session_id or parent_task.session_id or session_id or "").strip()
|
||||
v_frames, _v_src = collect_legacy_lane_images_with_session_fallback(
|
||||
store=self.store,
|
||||
session_id=policy_sid,
|
||||
attachments=list(parent_task.attachments or []),
|
||||
max_images=1,
|
||||
)
|
||||
v_frame = str(v_frames[0]).strip() if v_frames else None
|
||||
resp_v = send_video_generation_request(
|
||||
prompt=user_text_v,
|
||||
model=str(getattr(chosen_model, "model", "") or "").strip() or None,
|
||||
api_key=str(getattr(chosen_model, "api_key", "") or "").strip() or None,
|
||||
base_url=str(getattr(chosen_model, "base_url", "") or "").strip() or None,
|
||||
img_url=v_frame,
|
||||
on_progress=on_progress,
|
||||
should_stop=should_stop,
|
||||
)
|
||||
ok, output, produced_attachments = legacy_video_turn_bundle(resp_v)
|
||||
output = legacy_video_assistant_body_with_placeholder(
|
||||
lang=self.lang,
|
||||
body_text=output,
|
||||
produced=produced_attachments if ok else None,
|
||||
)
|
||||
self.store.add_message(
|
||||
session_id=session_id,
|
||||
role="assistant",
|
||||
content=output,
|
||||
attachments=(produced_attachments or None) if ok else None,
|
||||
)
|
||||
specialist_delivery = SpecialistDelivery(
|
||||
specialist=step.specialist,
|
||||
step_id=step.step_id,
|
||||
answer_text=str(output or ""),
|
||||
tool_traces=(),
|
||||
notes="video_pipeline",
|
||||
)
|
||||
else:
|
||||
agent = self._build_agent_for(
|
||||
step.specialist,
|
||||
|
|
|
|||
|
|
@ -43,10 +43,15 @@ SPECIALISTS: dict[SpecialistId, SpecialistConfig] = {
|
|||
expert_name="image",
|
||||
default_tool_tags=None,
|
||||
),
|
||||
"video": SpecialistConfig(
|
||||
specialist_id="video",
|
||||
expert_name="video",
|
||||
default_tool_tags=None,
|
||||
),
|
||||
}
|
||||
|
||||
def discover_specialist_ids() -> tuple[SpecialistId, ...]:
|
||||
rows = specialist_registry_snapshot(base_order=("generalist", "ops", "memory", "image"))
|
||||
rows = specialist_registry_snapshot(base_order=("generalist", "ops", "memory", "image", "video"))
|
||||
return tuple(str(x.get("id") or "").strip().lower() for x in rows if str(x.get("id") or "").strip())
|
||||
|
||||
|
||||
|
|
|
|||
|
|
@ -1206,6 +1206,97 @@ def _maybe_image_specialist_legacy_gateway_turn(
|
|||
)
|
||||
|
||||
|
||||
def _maybe_video_specialist_legacy_gateway_turn(
|
||||
*,
|
||||
store: Any,
|
||||
session_id: str,
|
||||
turn_uuid: str,
|
||||
lang: str,
|
||||
model: ChatModel,
|
||||
user_text: str,
|
||||
attachments: list[dict[str, Any]] | None,
|
||||
skill_binding_role: str | None,
|
||||
on_token: Optional[Callable[[str], None]],
|
||||
on_progress: Optional[Callable[[str], None]],
|
||||
should_stop: Optional[Callable[[], bool]] = None,
|
||||
) -> TurnRunOutcome | None:
|
||||
"""When the UI selects **video** specialist, skip Responses/chat-model transports.
|
||||
|
||||
Uses DashScope async ``video-synthesis`` (see :mod:`oclaw.platform.llm.video_generation_client`).
|
||||
With a user image (or session image fallback), sends ``input.img_url`` for **image-to-video**;
|
||||
otherwise **text-to-video**. Disable with ``AIA_VIDEO_SPECIALIST_DISABLE_LEGACY_GATEWAY_LANE=1``.
|
||||
"""
|
||||
if str(os.getenv("AIA_VIDEO_SPECIALIST_DISABLE_LEGACY_GATEWAY_LANE") or "").strip().lower() in (
|
||||
"1",
|
||||
"true",
|
||||
"yes",
|
||||
"on",
|
||||
):
|
||||
return None
|
||||
if str(skill_binding_role or "").strip().lower() != "video":
|
||||
return None
|
||||
|
||||
from oclaw.platform.llm.image_legacy_client import collect_legacy_lane_images_with_session_fallback
|
||||
from oclaw.platform.llm.video_generation_client import (
|
||||
VIDEO_SPECIALIST_DEFAULT_PROMPT_ZH,
|
||||
legacy_video_assistant_body_with_placeholder,
|
||||
legacy_video_turn_bundle,
|
||||
send_video_generation_request,
|
||||
)
|
||||
|
||||
frames, frame_src = collect_legacy_lane_images_with_session_fallback(
|
||||
store=store,
|
||||
session_id=session_id,
|
||||
attachments=attachments,
|
||||
max_images=1,
|
||||
)
|
||||
frame_url = str(frames[0]).strip() if frames else None
|
||||
|
||||
if on_progress:
|
||||
if frame_url:
|
||||
if frame_src.endswith("_history"):
|
||||
if str(lang or "").startswith("en"):
|
||||
on_progress("oclaw: reusing an earlier session image as first frame…")
|
||||
else:
|
||||
on_progress("oclaw: 使用会话中较早的图片作为图生视频首帧…")
|
||||
on_progress("oclaw: video specialist (DashScope image-to-video)…")
|
||||
else:
|
||||
on_progress("oclaw: video specialist (DashScope text-to-video)…")
|
||||
|
||||
prompt_plain = str(user_text or "").strip() or VIDEO_SPECIALIST_DEFAULT_PROMPT_ZH
|
||||
resp = send_video_generation_request(
|
||||
prompt=prompt_plain,
|
||||
model=str(getattr(model, "model", "") or "").strip() or None,
|
||||
api_key=str(getattr(model, "api_key", "") or "").strip() or None,
|
||||
base_url=str(getattr(model, "base_url", "") or "").strip() or None,
|
||||
img_url=frame_url,
|
||||
on_progress=on_progress,
|
||||
should_stop=should_stop,
|
||||
)
|
||||
ok, body_text, produced = legacy_video_turn_bundle(resp)
|
||||
body_text = legacy_video_assistant_body_with_placeholder(
|
||||
lang=lang,
|
||||
body_text=body_text,
|
||||
produced=produced if ok else None,
|
||||
)
|
||||
store.add_message(
|
||||
session_id=session_id,
|
||||
role="assistant",
|
||||
content=body_text,
|
||||
turn_uuid=turn_uuid,
|
||||
event_type="assistant_text",
|
||||
attachments=(produced or None) if ok else None,
|
||||
)
|
||||
if ok and on_token and body_text:
|
||||
on_token(body_text)
|
||||
return TurnRunOutcome(
|
||||
final_text=body_text,
|
||||
tool_traces=tuple(),
|
||||
handoff_note="video_specialist_legacy_http" if ok else "video_specialist_legacy_upstream_failed",
|
||||
turn_uuid=turn_uuid,
|
||||
)
|
||||
|
||||
|
||||
def run_oclaw_direct_loop(
|
||||
*,
|
||||
store: Any,
|
||||
|
|
@ -1267,6 +1358,21 @@ def run_oclaw_direct_loop(
|
|||
)
|
||||
if legacy_early is not None:
|
||||
return legacy_early
|
||||
video_early = _maybe_video_specialist_legacy_gateway_turn(
|
||||
store=store,
|
||||
session_id=session_id,
|
||||
turn_uuid=turn_uuid,
|
||||
lang=lang,
|
||||
model=model,
|
||||
user_text=str(user_text or ""),
|
||||
attachments=attachments,
|
||||
skill_binding_role=skill_binding_role,
|
||||
on_token=on_token,
|
||||
on_progress=on_progress,
|
||||
should_stop=should_stop,
|
||||
)
|
||||
if video_early is not None:
|
||||
return video_early
|
||||
|
||||
skill_exec = SkillExecutor(config=ToolExecutionConfig(max_workers=max(1, min(int(max_tool_workers or 8), 32))))
|
||||
tool_traces: list[dict[str, Any]] = []
|
||||
|
|
|
|||
|
|
@ -919,7 +919,7 @@ class OclawGateway:
|
|||
attachments=list(msg.attachments or []),
|
||||
metadata=dict(base_metadata),
|
||||
)
|
||||
if manager_specialist in {"ops", "generalist", "image", "memory"}:
|
||||
if manager_specialist in {"ops", "generalist", "image", "memory", "video"}:
|
||||
if callable(specialist_executor_factory):
|
||||
try:
|
||||
selected_executor = specialist_executor_factory(manager_specialist)
|
||||
|
|
|
|||
|
|
@ -17,7 +17,7 @@ def _truthy(v: str | None) -> bool:
|
|||
|
||||
def ordered_specialist_ids() -> list[str]:
|
||||
base = [str(k).strip().lower() for k in discover_specialist_ids() if str(k).strip()]
|
||||
preferred = [x for x in ("generalist", "ops", "memory", "image") if x in set(base)]
|
||||
preferred = [x for x in ("generalist", "ops", "memory", "image", "video") if x in set(base)]
|
||||
return preferred + [x for x in base if x not in set(preferred)]
|
||||
|
||||
|
||||
|
|
|
|||
|
|
@ -252,7 +252,7 @@ def build_expert_catalog_block(*, include_main: bool = False, per_field_limit: i
|
|||
|
||||
def discover_specialist_ids_from_workspaces(
|
||||
*,
|
||||
base_order: tuple[str, ...] = ("generalist", "ops", "memory", "image"),
|
||||
base_order: tuple[str, ...] = ("generalist", "ops", "memory", "image", "video"),
|
||||
) -> tuple[str, ...]:
|
||||
cache_key = (expert_workspace_signature_token(), tuple(str(x).strip().lower() for x in base_order if str(x).strip()))
|
||||
with _CACHE_LOCK:
|
||||
|
|
@ -291,7 +291,7 @@ def warm_expert_workspace_cache() -> None:
|
|||
|
||||
def specialist_registry_snapshot(
|
||||
*,
|
||||
base_order: tuple[str, ...] = ("generalist", "ops", "memory", "image"),
|
||||
base_order: tuple[str, ...] = ("generalist", "ops", "memory", "image", "video"),
|
||||
) -> tuple[dict[str, Any], ...]:
|
||||
"""Single source of truth for runtime specialist discovery and metadata."""
|
||||
ordered = discover_specialist_ids_from_workspaces(base_order=base_order)
|
||||
|
|
|
|||
5
runtime/workspaces/video/META.json
Normal file
5
runtime/workspaces/video/META.json
Normal file
|
|
@ -0,0 +1,5 @@
|
|||
{
|
||||
"display_name_en": "Video generation",
|
||||
"display_name_zh": "视频生成专家",
|
||||
"role": "expert"
|
||||
}
|
||||
1
runtime/workspaces/video/ROLE_SYSTEM.md
Normal file
1
runtime/workspaces/video/ROLE_SYSTEM.md
Normal file
|
|
@ -0,0 +1 @@
|
|||
你是视频生成专家:将用户自然语言 prompt 交给 Wan / 百炼 text-to-video API,返回可下载或可播放的成片附件。不要编造已生成视频的 URL;仅展示接口真实返回结果或明确错误。
|
||||
9
runtime/workspaces/video/SOUL.md
Normal file
9
runtime/workspaces/video/SOUL.md
Normal file
|
|
@ -0,0 +1,9 @@
|
|||
你是**文生视频 / 图生视频**方向的专家助手:根据用户给出的画面与镜头描述(及可选的**首帧参考图**),调用百炼 / DashScope **异步视频合成**接口生成短视频结果,并将产出以会话附件(`video_ref`)形式返回。用户上传图片时,首帧会作为 `img_url` 提交(需使用支持图生视频的 i2v 模型)。
|
||||
|
||||
回答要求:
|
||||
- 若用户描述含糊,可基于常识补全合理的镜头语言,但避免与用户明确约束相矛盾。
|
||||
- 生成失败时给出可读的上游错误或参数提示(如模型与区域、时长、分辨率不匹配)。
|
||||
|
||||
边界:
|
||||
- 本专家链路**不调用**通用工具循环;仅走专用 HTTP 视频合成与轮询。
|
||||
- 不承诺具体成片内容符合版权素材或真人肖像等合规要求;用户需自行确保 prompt 合规。
|
||||
27
tests/test_chat_attachments_payload.py
Normal file
27
tests/test_chat_attachments_payload.py
Normal file
|
|
@ -0,0 +1,27 @@
|
|||
"""Regression: admin REST and WS share attachment normalization (typed refs must survive parsing)."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from oclaw.interfaces.admin.chat_api import _parse_attachments_payload
|
||||
|
||||
|
||||
def test_parse_attachments_payload_preserves_image_ref() -> None:
|
||||
att: dict = {"type": "image_ref", "attachment_id": "a" * 64, "mime": "image/png"}
|
||||
out = _parse_attachments_payload([att])
|
||||
assert out == [att]
|
||||
|
||||
|
||||
def test_parse_attachments_payload_preserves_relay_pointer_image() -> None:
|
||||
att: dict = {
|
||||
"type": "relay_pointer",
|
||||
"mime": "image/jpeg",
|
||||
"attachment_id": "b" * 64,
|
||||
"pointer_uri": "relay://attachments/x/" + "b" * 64,
|
||||
}
|
||||
out = _parse_attachments_payload([att])
|
||||
assert out == [att]
|
||||
|
||||
|
||||
def test_parse_attachments_payload_empty() -> None:
|
||||
assert _parse_attachments_payload(None) is None
|
||||
assert _parse_attachments_payload([]) is None
|
||||
|
|
@ -575,3 +575,16 @@ def test_excel_sheet_count_cap_applies_in_tool_mode(tmp_path: Path, monkeypatch)
|
|||
notes = [x for x in out if isinstance(x, dict) and str(x.get("name") or "").endswith(".sheet-limit")]
|
||||
assert notes
|
||||
|
||||
|
||||
def test_attachment_asset_store_mp4_load_bytes_roundtrip(tmp_path: Path) -> None:
|
||||
"""Regression: ``load_bytes`` must resolve ``.mp4`` on-disk blobs (not only image extensions)."""
|
||||
from oclaw.platform.files.attachment_assets import AttachmentAssetStore
|
||||
|
||||
ast = AttachmentAssetStore(str(tmp_path))
|
||||
payload = b"\x00\x00\x00\x18ftypmp42\x00\x00\x00\x00mp42isom" + b"vid" * 400
|
||||
meta = ast.save_bytes(payload, filename="clip.mp4", mime="video/mp4")
|
||||
blob, meta2 = ast.load_bytes(meta.attachment_id)
|
||||
assert blob == payload
|
||||
assert meta2 is not None
|
||||
assert meta2.bytes == len(payload)
|
||||
|
||||
|
|
|
|||
257
tests/test_video_generation_client.py
Normal file
257
tests/test_video_generation_client.py
Normal file
|
|
@ -0,0 +1,257 @@
|
|||
from __future__ import annotations
|
||||
|
||||
from typing import Any
|
||||
from unittest.mock import MagicMock, patch
|
||||
|
||||
from oclaw.platform.llm.video_generation_client import (
|
||||
_effective_video_model_for_request,
|
||||
_i2v_first_frame_input_style,
|
||||
dashscope_api_root_from_base_url,
|
||||
legacy_video_assistant_body_with_placeholder,
|
||||
legacy_video_turn_bundle,
|
||||
materialize_video_output_attachments,
|
||||
send_video_generation_request,
|
||||
)
|
||||
|
||||
|
||||
def test_dashscope_api_root_strips_compatible_suffix() -> None:
|
||||
assert (
|
||||
dashscope_api_root_from_base_url("https://dashscope.aliyuncs.com/compatible-mode/v1")
|
||||
== "https://dashscope.aliyuncs.com"
|
||||
)
|
||||
|
||||
|
||||
def test_send_video_generation_missing_config() -> None:
|
||||
r = send_video_generation_request(prompt="hello", api_key="", base_url="", model="wan2.2-t2v-plus")
|
||||
assert r["ok"] is False
|
||||
assert "missing" in str(r.get("error") or "").lower()
|
||||
|
||||
|
||||
def test_send_video_generation_poll_succeeds() -> None:
|
||||
post_resp = MagicMock()
|
||||
post_resp.status_code = 200
|
||||
post_resp.json.return_value = {"output": {"task_id": "tid-1", "task_status": "PENDING"}}
|
||||
|
||||
get_pending = MagicMock()
|
||||
get_pending.status_code = 200
|
||||
get_pending.json.return_value = {"output": {"task_id": "tid-1", "task_status": "PENDING"}}
|
||||
|
||||
get_ok = MagicMock()
|
||||
get_ok.status_code = 200
|
||||
get_ok.json.return_value = {
|
||||
"output": {
|
||||
"task_id": "tid-1",
|
||||
"task_status": "SUCCEEDED",
|
||||
"video_url": "https://example.invalid/out.mp4",
|
||||
}
|
||||
}
|
||||
|
||||
clients: list[MagicMock] = []
|
||||
|
||||
def _client_factory(*_a, **_k):
|
||||
inst = MagicMock()
|
||||
inst.__enter__ = lambda *_x: inst
|
||||
inst.__exit__ = lambda *_x: None
|
||||
if len(clients) == 1:
|
||||
inst.get.side_effect = [get_pending, get_ok]
|
||||
clients.append(inst)
|
||||
return inst
|
||||
|
||||
with patch("oclaw.platform.llm.video_generation_client.httpx.Client", side_effect=_client_factory):
|
||||
with patch(
|
||||
"oclaw.platform.llm.video_generation_client.post_with_retry",
|
||||
return_value=post_resp,
|
||||
):
|
||||
with patch("oclaw.platform.llm.video_generation_client.time.sleep"):
|
||||
r = send_video_generation_request(
|
||||
prompt="cat",
|
||||
model="wan2.2-t2v-plus",
|
||||
api_key="sk-test",
|
||||
base_url="https://dashscope.aliyuncs.com",
|
||||
)
|
||||
|
||||
assert r["ok"] is True
|
||||
assert r.get("video_urls") == ["https://example.invalid/out.mp4"]
|
||||
assert len(clients) == 2
|
||||
assert clients[1].get.call_count == 2
|
||||
|
||||
|
||||
def test_legacy_video_turn_bundle_ok_without_materialize() -> None:
|
||||
ok, text, att = legacy_video_turn_bundle(
|
||||
{"ok": True, "text": "done", "video_urls": ["https://example.invalid/x.mp4"]}
|
||||
)
|
||||
assert ok is True
|
||||
assert text == "done"
|
||||
assert len(att) >= 1
|
||||
assert att[0].get("type") == "video_ref"
|
||||
|
||||
|
||||
def test_legacy_video_placeholder_zh() -> None:
|
||||
produced = [{"type": "video_ref", "attachment_id": "a" * 64}]
|
||||
s = legacy_video_assistant_body_with_placeholder(lang="zh", body_text="", produced=produced)
|
||||
assert "视频" in s
|
||||
|
||||
|
||||
def test_materialize_video_output_attachments_empty() -> None:
|
||||
assert materialize_video_output_attachments([], max_videos=1) == []
|
||||
|
||||
|
||||
def test_effective_video_model_coerces_t2v_to_i2v_when_image() -> None:
|
||||
assert _effective_video_model_for_request("wan2.6-t2v-flash", img_url="https://x/a.png") == "wan2.6-i2v-flash"
|
||||
assert _effective_video_model_for_request("wan2.6-i2v-flash", img_url="https://x/a.png") == "wan2.6-i2v-flash"
|
||||
assert _effective_video_model_for_request("wan2.6-t2v-flash", img_url=None) == "wan2.6-t2v-flash"
|
||||
assert _effective_video_model_for_request("wan2.6-t2v-flash", img_url="") == "wan2.6-t2v-flash"
|
||||
assert _effective_video_model_for_request("wan2.7-t2v", img_url="https://x/a.png") == "wan2.7-i2v"
|
||||
|
||||
|
||||
def test_i2v_input_style_auto(monkeypatch) -> None:
|
||||
monkeypatch.delenv("AIA_VIDEO_I2V_INPUT_STYLE", raising=False)
|
||||
assert _i2v_first_frame_input_style(model_name="wan2.7-i2v") == "media"
|
||||
assert _i2v_first_frame_input_style(model_name="WAN2.7-i2v-plus") == "media"
|
||||
assert _i2v_first_frame_input_style(model_name="wan2.6-i2v-flash") == "img_url"
|
||||
|
||||
|
||||
def test_send_video_coerces_payload_model_when_t2v_plus_img(monkeypatch) -> None:
|
||||
monkeypatch.delenv("AIA_VIDEO_EXPERT_I2V_MODEL", raising=False)
|
||||
monkeypatch.delenv("AIA_VIDEO_EXPERT_DISABLE_I2V_MODEL_COERCION", raising=False)
|
||||
post_resp = MagicMock()
|
||||
post_resp.status_code = 200
|
||||
post_resp.json.return_value = {
|
||||
"output": {
|
||||
"task_id": "tid-i2v",
|
||||
"task_status": "SUCCEEDED",
|
||||
"video_url": "https://example.invalid/i2v.mp4",
|
||||
}
|
||||
}
|
||||
inst = MagicMock()
|
||||
inst.__enter__ = lambda *_x: inst
|
||||
inst.__exit__ = lambda *_x: None
|
||||
captured: dict[str, Any] = {}
|
||||
|
||||
def _capture_post(_client, **kw):
|
||||
p = kw.get("payload")
|
||||
if isinstance(p, dict):
|
||||
captured.clear()
|
||||
captured.update(p)
|
||||
return post_resp
|
||||
|
||||
with patch("oclaw.platform.llm.video_generation_client.httpx.Client", return_value=inst):
|
||||
with patch("oclaw.platform.llm.video_generation_client.post_with_retry", side_effect=_capture_post):
|
||||
r = send_video_generation_request(
|
||||
prompt="motion",
|
||||
model="wan2.6-t2v-flash",
|
||||
api_key="k",
|
||||
base_url="https://dashscope.aliyuncs.com",
|
||||
img_url="https://example.invalid/frame.png",
|
||||
)
|
||||
assert r["ok"] is True
|
||||
assert captured.get("model") == "wan2.6-i2v-flash"
|
||||
|
||||
|
||||
def test_send_video_includes_img_url_in_payload() -> None:
|
||||
post_resp = MagicMock()
|
||||
post_resp.status_code = 200
|
||||
post_resp.json.return_value = {
|
||||
"output": {
|
||||
"task_id": "tid-i2v",
|
||||
"task_status": "SUCCEEDED",
|
||||
"video_url": "https://example.invalid/i2v.mp4",
|
||||
}
|
||||
}
|
||||
inst = MagicMock()
|
||||
inst.__enter__ = lambda *_x: inst
|
||||
inst.__exit__ = lambda *_x: None
|
||||
captured: dict[str, Any] = {}
|
||||
|
||||
def _capture_post(_client, **kw):
|
||||
p = kw.get("payload")
|
||||
if isinstance(p, dict):
|
||||
captured.clear()
|
||||
captured.update(p)
|
||||
return post_resp
|
||||
|
||||
with patch("oclaw.platform.llm.video_generation_client.httpx.Client", return_value=inst):
|
||||
with patch("oclaw.platform.llm.video_generation_client.post_with_retry", side_effect=_capture_post):
|
||||
r = send_video_generation_request(
|
||||
prompt="motion",
|
||||
model="wan2.6-i2v-flash",
|
||||
api_key="k",
|
||||
base_url="https://dashscope.aliyuncs.com",
|
||||
img_url="https://example.invalid/frame.png",
|
||||
)
|
||||
assert r["ok"] is True
|
||||
inp = captured.get("input")
|
||||
assert isinstance(inp, dict)
|
||||
assert inp.get("img_url") == "https://example.invalid/frame.png"
|
||||
assert inp.get("media") is None
|
||||
|
||||
|
||||
def test_send_video_wan27_uses_media_first_frame(monkeypatch) -> None:
|
||||
monkeypatch.delenv("AIA_VIDEO_I2V_INPUT_STYLE", raising=False)
|
||||
post_resp = MagicMock()
|
||||
post_resp.status_code = 200
|
||||
post_resp.json.return_value = {
|
||||
"output": {
|
||||
"task_id": "tid-i2v",
|
||||
"task_status": "SUCCEEDED",
|
||||
"video_url": "https://example.invalid/i2v.mp4",
|
||||
}
|
||||
}
|
||||
inst = MagicMock()
|
||||
inst.__enter__ = lambda *_x: inst
|
||||
inst.__exit__ = lambda *_x: None
|
||||
captured: dict[str, Any] = {}
|
||||
|
||||
def _capture_post(_client, **kw):
|
||||
p = kw.get("payload")
|
||||
if isinstance(p, dict):
|
||||
captured.clear()
|
||||
captured.update(p)
|
||||
return post_resp
|
||||
|
||||
u = "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20250925/wpimhv/rap.png"
|
||||
with patch("oclaw.platform.llm.video_generation_client.httpx.Client", return_value=inst):
|
||||
with patch("oclaw.platform.llm.video_generation_client.post_with_retry", side_effect=_capture_post):
|
||||
r = send_video_generation_request(
|
||||
prompt="你好",
|
||||
model="wan2.7-i2v",
|
||||
api_key="k",
|
||||
base_url="https://dashscope.aliyuncs.com",
|
||||
img_url=u,
|
||||
)
|
||||
assert r["ok"] is True
|
||||
inp = captured.get("input")
|
||||
assert isinstance(inp, dict)
|
||||
assert inp.get("img_url") is None
|
||||
media = inp.get("media")
|
||||
assert isinstance(media, list) and len(media) >= 1
|
||||
assert media[0].get("type") == "first_frame"
|
||||
assert media[0].get("url") == u
|
||||
|
||||
|
||||
def test_send_immediate_succeeded_on_submit() -> None:
|
||||
post_resp = MagicMock()
|
||||
post_resp.status_code = 200
|
||||
post_resp.json.return_value = {
|
||||
"output": {
|
||||
"task_id": "tid-x",
|
||||
"task_status": "SUCCEEDED",
|
||||
"video_url": "https://example.invalid/immediate.mp4",
|
||||
}
|
||||
}
|
||||
inst = MagicMock()
|
||||
inst.__enter__ = lambda *_x: inst
|
||||
inst.__exit__ = lambda *_x: None
|
||||
with patch("oclaw.platform.llm.video_generation_client.httpx.Client", return_value=inst):
|
||||
with patch(
|
||||
"oclaw.platform.llm.video_generation_client.post_with_retry",
|
||||
return_value=post_resp,
|
||||
):
|
||||
r = send_video_generation_request(
|
||||
prompt="x",
|
||||
model="m",
|
||||
api_key="k",
|
||||
base_url="https://dashscope.aliyuncs.com",
|
||||
)
|
||||
assert r["ok"] is True
|
||||
assert r["video_urls"] == ["https://example.invalid/immediate.mp4"]
|
||||
Loading…
Add table
Add a link
Reference in a new issue