Hermes + Image to video
Hermes runs the work loop and Image to video performs the selected operation. The Composer below preselects Hermes, DeepSeek V4 Pro, and this tool so you can test one real task with the same billing and recovery path as any AgentSky run.
How the combination works
Hermes decides when to call Image to video.
Hermes coordinates a general-purpose agent loop across research, files, shell commands, and the tools enabled for the task.
Start with a publicly reachable image URL and describe the motion you want. The current adapter defaults to Wan 2.6 I2V Flash through DashScope and requests a 1280×720 video. The historical mm.i2v identifier does not indicate a MiniMax model.
Hermes
Owns the task loop, context, model calls, tool choice, cloud computer, progress and final result.
Image to video
Performs the operation exposed by the current Image to video capability. Its input limits and output shape come from the live tool guide and adapter.
AgentSky
Keeps the selected agent, model, capability grant, task state and usage record together. Tool charges and model or runtime charges remain itemized separately.
Setup and first task
Start with one bounded image to video task.
The Composer carries Hermes, DeepSeek V4 Pro, and Image to video. Begin with the smallest input that lets you inspect the output before expanding the workflow.
Open the Hermes Composer
The Composer above preselects Hermes, DeepSeek V4 Pro, and Image to video. Add the target, required inputs and a concrete expected result; the draft stays in place through sign-in.
State the Image to video boundary
Animate this landscape with a gentle camera push and subtle movement in the trees. Preserve the central composition.
Run and verify the returned result
Check the output against the source or brief, confirm that the result has the expected format, and inspect the task usage for model, runtime and tool charges.
Inputs, cost and limits
What Hermes + Image to video establishes
Research and operational work that needs a flexible tool-using agent rather than a repository-specific coding interface.
Configure and verify
- Hermes and DeepSeek V4 Pro are preselected for the browser evaluation.
- Image to video is enabled from the current capability catalog with its live input and output guidance.
- The first task can be checked from the returned artifact, text or provider result before a larger workflow is attempted.
Still depends on the task
- A tool guide does not guarantee a provider result for malformed, blocked or unsupported Image to video inputs.
- A successful tool call does not establish that the result is correct for your business decision; inspect the output at its source boundary.
- Total task cost can include model tokens, active computer time and the selected tool's billable units.
Current Image to video rates
mm.i2v
$0.0488 per second.
Tool-specific details
Use the current Image to video contract.
Start with a publicly reachable image URL and describe the motion you want. The current adapter defaults to Wan 2.6 I2V Flash through DashScope and requests a 1280×720 video. The historical mm.i2v identifier does not indicate a MiniMax model.
Required input
A public image URL and a motion prompt. The current execution path rejects a local file path or private URL it cannot retrieve.
Default model and output
wan2.6-i2v-flash, configurable by deployment; 1280×720 output, with prompt extension disabled.
Duration and delivery
Duration defaults to five seconds. The catalog accepts 1–10; the provider must also support the chosen value. The completed job returns an MP4 artifact.
Common questions
Related resources
