feat(video): DashScope video specialist with correct i2v request bodies

- Add video_generation_client: async video-synthesis, t2v vs i2v (Wan 2.7 input.media first_frame vs legacy img_url).

- Coerce *-t2v* to *-i2v* when first frame present; AIA_VIDEO_I2V_INPUT_STYLE / I2V_MODEL overrides.

- Gateway/direct_loop early return for video specialist; workspace video + factory allowlist.

- Normalize WS and admin chat attachments (image_ref parity); mp4 attachment store roundtrip tests.

Co-authored-by: Cursor <cursoragent@cursor.com>
This commit is contained in:
oliver 2026-05-10 16:20:43 +08:00
parent 398d85ba06
commit 4335040a27
21 changed files with 1152 additions and 56 deletions

View file

@ -252,7 +252,7 @@ def build_expert_catalog_block(*, include_main: bool = False, per_field_limit: i
def discover_specialist_ids_from_workspaces(
*,
base_order: tuple[str, ...] = ("generalist", "ops", "memory", "image"),
base_order: tuple[str, ...] = ("generalist", "ops", "memory", "image", "video"),
) -> tuple[str, ...]:
cache_key = (expert_workspace_signature_token(), tuple(str(x).strip().lower() for x in base_order if str(x).strip()))
with _CACHE_LOCK:
@ -291,7 +291,7 @@ def warm_expert_workspace_cache() -> None:
def specialist_registry_snapshot(
*,
base_order: tuple[str, ...] = ("generalist", "ops", "memory", "image"),
base_order: tuple[str, ...] = ("generalist", "ops", "memory", "image", "video"),
) -> tuple[dict[str, Any], ...]:
"""Single source of truth for runtime specialist discovery and metadata."""
ordered = discover_specialist_ids_from_workspaces(base_order=base_order)

View file

@ -0,0 +1,5 @@
{
"display_name_en": "Video generation",
"display_name_zh": "视频生成专家",
"role": "expert"
}

View file

@ -0,0 +1 @@
你是视频生成专家:将用户自然语言 prompt 交给 Wan / 百炼 text-to-video API,返回可下载或可播放的成片附件。不要编造已生成视频的 URL;仅展示接口真实返回结果或明确错误。

View file

@ -0,0 +1,9 @@
你是**文生视频 / 图生视频**方向的专家助手:根据用户给出的画面与镜头描述(及可选的**首帧参考图**),调用百炼 / DashScope **异步视频合成**接口生成短视频结果,并将产出以会话附件(`video_ref`)形式返回。用户上传图片时,首帧会作为 `img_url` 提交(需使用支持图生视频的 i2v 模型)。
回答要求:
- 若用户描述含糊,可基于常识补全合理的镜头语言,但避免与用户明确约束相矛盾。
- 生成失败时给出可读的上游错误或参数提示(如模型与区域、时长、分辨率不匹配)。
边界:
- 本专家链路**不调用**通用工具循环;仅走专用 HTTP 视频合成与轮询。
- 不承诺具体成片内容符合版权素材或真人肖像等合规要求;用户需自行确保 prompt 合规。