mirror of
https://github.com/hansjone/oclaw.git
synced 2026-10-09 07:20:44 +08:00
- Add video_generation_client: async video-synthesis, t2v vs i2v (Wan 2.7 input.media first_frame vs legacy img_url). - Coerce *-t2v* to *-i2v* when first frame present; AIA_VIDEO_I2V_INPUT_STYLE / I2V_MODEL overrides. - Gateway/direct_loop early return for video specialist; workspace video + factory allowlist. - Normalize WS and admin chat attachments (image_ref parity); mp4 attachment store roundtrip tests. Co-authored-by: Cursor <cursoragent@cursor.com>
4.7 KiB
4.7 KiB
LLM provider/transport capability matrix (oclaw)
This project follows an Oclaw-style explicit provider/transport selection:
- You select the transport via LLM profile
mode(not by inferring frombase_url). base_urlcan be the same unified gateway URL for all providers; themodedetermines the wire protocol.
Profile field mapping
mode: transport selector (openai,openai_responses,anthropic,google,ollama,rule)model: model id/name passed to the provider transportbase_url: gateway host URL (may be shared across providers)api_key(profile secret): primary credential source (reused across modes)
Transport selection happens in oclaw/runtime/agents/factory.py.
Modes and transports
openai (Chat Completions, OpenAI-compatible)
- Transport:
oclaw/platform/llm/transports/openai_chat_completions.py::OpenAIChatModel - API:
/v1/chat/completions(streaming supported) - Tools: OpenAI tool calling (
tools[]/tool_calls) - Streaming: token deltas via
on_token→ WSchat.delta - Key: profile secret or
OPENAI_API_KEY
openai_responses (Responses API, OpenAI-compatible)
- Transport:
oclaw/platform/llm/transports/openai_responses.py::OpenAIResponsesModel - API:
/v1/responses(streaming events) - Tools: function call items parsed from response output
- Streaming: output text deltas via
on_token→ WSchat.delta - Key: profile secret or
OPENAI_API_KEY
Chat UI「图片专家」(绕行本矩阵)
- Not a separate profile transport: when the user selects specialist
imagein/chat,runtime/direct_loop.pytakes an early return and callsplatform/llm/image_legacy_client.send_legacy_image_messages(DashScope-style/chat/completions), so that turn does not useOpenAIResponsesModel/ chat transports above. - Details, env vars, ACL, and UI hooks:
docs/IMAGE_SPECIALIST_LANE.md.
Chat UI「视频生成专家」(绕行本矩阵)
- When the user selects specialist
video,runtime/direct_loop.pyearly-returns intoplatform/llm/video_generation_client.send_video_generation_request(DashScope asyncvideo-synthesis+ task polling). Without an input image: text-to-video; with an image attachment (or session image fallback):input.img_urlfor image-to-video (use an i2v model id). That turn does not use the generic tool loop orOpenAIResponsesModel. docs/VIDEO_SPECIALIST_LANE.md.
anthropic (Anthropic Messages streaming)
- Transport:
oclaw/platform/llm/transports/anthropic_messages.py::AnthropicMessagesModel - API: Anthropic
messages.streamsurface (gateway must provide Anthropic-compatible protocol) - Tools: tool use blocks assembled into
LLMToolCall - Streaming: text deltas via
on_token→ WSchat.delta - Key: profile secret or
ANTHROPIC_API_KEY(fallback:OPENAI_API_KEYfor unified gateways)
google (Gemini native SSE)
- Transport:
oclaw/platform/llm/transports/google_gemini_sse.py::GoogleGeminiChatModel - API:
:streamGenerateContent?alt=sse(Gemini native) - Tools:
functionDeclarationswithparametersJsonSchema; parsesfunctionCall - Streaming: text deltas via
on_token→ WSchat.delta - Key: profile secret or
GOOGLE_API_KEY/GEMINI_API_KEY(fallback:OPENAI_API_KEYfor unified gateways) - Thinking controls (optional env):
AIA_GEMINI_THINKING=on|offAIA_GEMINI_THINKING_LEVEL=<string>AIA_GEMINI_THINKING_BUDGET=<int>
ollama (local OpenAI-compatible)
- Transport:
OpenAIChatModelwith Ollama-compatible base url - API:
/v1/chat/completions(Ollama OpenAI-compat) - Key: uses a dummy key (
ollama) if needed by SDK
rule (no remote LLM)
- Transport:
oclaw/platform/llm/transports/simple.py::RuleBasedChatModel - Purpose: deterministic tool routing/diagnostics without any model provider
Streaming and UI contract
All transports stream assistant output through the same internal callback:
- Transport calls
on_token(text_delta) - WS gateway maps that to
event=chat payload.state=delta - Final persisted assistant message is emitted as:
event=chat payload.state=final(withmessagepayload for folding)event=session.message
- Tool UI signals from runtime are emitted as
event=session.tool
Adding the remaining models
For additional providers, follow this pattern:
- Add a new transport class under
oclaw/platform/llm/transports/ - Extend
modeselection inoclaw/runtime/agents/factory.py - Add an offline stream parser test under
tests/ - Add/verify WS contract tests (delta + final + session.tool)