top of page

OpenRouter for Image and Video Models: Multimodal Access, Provider Choice, Creative Workflows, Pricing, and Production Controls

  • 3 hours ago
  • 29 min read

OpenRouter has expanded beyond text-model routing into a broader multimodal layer through which developers can analyse images and videos, generate still images, create video clips, compare providers, and manage several creative-model families through one account and API environment.

The platform does not provide one proprietary image or video model of its own, because its role is to connect applications with models from companies such as OpenAI, Google, ByteDance, Black Forest Labs, Recraft, xAI, Krea, Kuaishou, Alibaba, MiniMax, and other providers.

This structure allows a creative application to move between fast concept models, detailed image generators, reference-based editors, vector systems, and production-oriented video models without building a completely separate integration and billing relationship for every provider.

The common OpenRouter layer does not make all models technically interchangeable, since image and video endpoints differ in reference limits, aspect ratios, resolutions, output formats, editing controls, audio support, pricing units, safety rules, retention policies, and generation time.

A reliable workflow must therefore distinguish media understanding from media creation, model capabilities from provider-endpoint capabilities, and low-cost experimentation from the controlled configuration used for final production.

The practical value of OpenRouter appears when an organization uses the platform as a discovery, testing, routing, billing, and orchestration system through which several models can serve different stages of one creative process.

·····

Multimodal access includes understanding, image creation, and video creation as separate operations.

A model that accepts an image can inspect a photograph, screenshot, chart, interface, drawing, or scanned page and return a written answer, although that capability does not automatically allow the same model to create a new visual asset.

A video-understanding model can describe scenes, identify actions, extract information, or answer questions about a clip, while video generation requires a different endpoint and a model designed to synthesize moving images.

OpenRouter separates these workflows through chat-based multimodal inputs, a dedicated synchronous Image API, and a dedicated asynchronous Video API.

This distinction matters because an application may use one model to analyse a reference image, another to generate a revised still, a third to animate the approved frame, and a fourth to review the completed output.

The word multimodal should therefore be interpreted as support for several forms of input or output rather than as evidence that every model performs every visual task.

........

The principal multimodal workflows available through OpenRouter.

Workflow

Input

Output

Typical OpenRouter interface

Image understanding

Image with optional text instructions

Text

Chat Completions

Video understanding

Video with optional text instructions

Text

Chat Completions

Image generation

Text with optional reference images

Image

Image API

Image editing

Existing image with written transformation instructions

Revised image

Image API

Video generation

Text with optional frames or references

Video

Video API

Agent-directed image creation

Natural-language objective

Text response and generated image

Image-generation server tool

Automated media review

Generated image or video

Written evaluation or structured result

Vision-capable chat model

·····

Image and video understanding remain useful before generation begins.

A creative process often starts with existing material, including a product photograph, campaign reference, design board, storyboard, previous advertisement, interface screenshot, or competitor example.

A vision-capable model can identify visual elements, describe composition, extract visible text, explain the apparent lighting, identify colour relationships, and convert an informal visual reference into a structured creative brief.

Video understanding can serve a similar role by identifying scenes, camera movement, objects, spoken content, transitions, visual continuity, and the sequence of actions within a clip.

This analytical stage can reduce prompt ambiguity because the generation model receives a more precise description of the approved reference rather than a vague instruction such as “make something like this.”

The model’s description should still be reviewed, particularly when exact text, small product details, identities, measurements, regulated claims, or subtle brand features matter.

Visual analysis creates a useful bridge between human references and generation prompts, although it does not replace the original media as the authoritative source.

........

Tasks suited to multimodal understanding models.

Input material

Possible analysis

Product photograph

Object placement, materials, reflections, and lighting

Advertising poster

Typography, hierarchy, colours, and composition

Interface screenshot

Components, layout, spacing, and visible states

Character reference sheet

Clothing, facial traits, accessories, and proportions

Storyboard

Shot order, framing, movement, and scene transitions

Existing video

Actions, camera behaviour, audio, and continuity

Chart or infographic

Labels, values, visual relationships, and possible errors

Scanned document

Text extraction, tables, signatures, and layout interpretation

·····

The dedicated Image API normalizes common generation controls.

OpenRouter’s Image API accepts a model identifier, a written prompt, and optional parameters that describe the requested asset.

Compatible endpoints may support square, portrait, landscape, ultrawide, and tall aspect ratios, while some models allow the application to request an exact pixel size instead of a predefined shape.

Resolution controls may extend from smaller concept images to 2K or 4K output, although the available levels depend on the selected model and provider endpoint.

The API can expose PNG, JPEG, WebP, and, for compatible vector models, SVG output.

Some endpoints support transparent backgrounds, compression settings, seeds, multiple images in one request, reference images, quality levels, and streamed partial renders.

The normalized structure reduces integration work, but an application must still inspect the model and endpoint documentation before exposing a control to users.

A parameter accepted by one image model may be ignored, transformed, or rejected by another.

........

Common image-generation controls available through compatible endpoints.

Control

Creative purpose

Prompt

Defines subject, environment, composition, lighting, and style

Aspect ratio

Adapts output to social, presentation, web, print, or mobile formats

Exact size

Requests specific width and height

Resolution

Balances detail, generation time, and cost

Quality

Selects a faster draft or higher-detail render

Output format

Produces PNG, JPEG, WebP, or SVG where supported

Background

Requests opaque or transparent output

Compression

Reduces file size for lossy formats

Seed

Supports limited repeatability where available

Number of images

Creates several variations in one request

Input references

Guides identity, product, composition, or style

Streaming

Returns partial visual progress before completion

·····

Model-level capabilities may combine features that no single endpoint supports together.

OpenRouter may expose one image model through several providers.

The model-level page can display the union of the parameters available across all active endpoints, which means that a listed capability is not necessarily supported by every provider serving that model.

One provider may support transparent backgrounds but not streaming, while another may offer streaming but restrict resolution or reference-image count.

The per-endpoint record should therefore be treated as the authoritative source for production configuration.

Applications that rely only on the model-level summary may present controls that fail when the router selects an endpoint whose implementation does not include them.

A robust interface can retrieve the available endpoints, filter them according to the required parameters, and then pin or rank the remaining providers.

........

Endpoint details that should be checked before sending a generation request.

Endpoint field

Operational importance

Provider identity

Shows which company receives and serves the request

Supported parameters

Confirms the definitive controls accepted by the endpoint

Passthrough parameters

Reveals provider-specific options

Resolution variants

Shows which dimensions or quality levels have separate pricing

Reference limit

Determines how many guidance images may be supplied

Streaming support

Determines whether partial images can be returned

Pricing unit

Identifies per-image, per-megapixel, or token billing

Data policy

Determines retention and training treatment

Availability

Shows whether the endpoint is active, preview, free, or restricted

Provider tag

Allows a tested endpoint to be selected directly

·····

OpenRouter’s image catalogue covers several distinct creative roles.

Image models should not be compared through one broad category called quality, because a photorealistic campaign image, a vector icon, a product edit, and a rapid concept sketch require different capabilities.

OpenAI’s GPT Image family is positioned around detailed generation, prompt adherence, text handling, and editing.

Google’s Gemini image models combine multimodal reasoning with generation and may be useful for complex composition, visual text, and image transformation.

ByteDance Seedream models emphasize visual quality, portraits, text, editing, and multi-reference composition.

Black Forest Labs offers several FLUX variants that cover economical generation, higher-quality rendering, and reference-driven workflows.

Recraft provides models designed for professional design outputs, including vector graphics and scalable SVG assets.

xAI’s Grok Imagine models target photorealistic and branded image creation, while Krea and other fast-generation providers focus on rapid experimentation and interactive ideation.

The appropriate shortlist begins with the required asset type rather than the model’s general popularity.

........

Representative image-model roles available through OpenRouter.

Creative requirement

Model categories worth evaluating

Fast concept exploration

Krea Turbo, FLUX Klein, and other low-cost rapid models

Detailed photographic output

GPT Image, Grok Imagine, Seedream, and higher FLUX variants

Text-heavy advertisements

GPT Image, Gemini image models, and text-capable visual systems

Product-preserving edits

Models with strong reference and editing support

Character consistency

Models accepting several identity or style references

Transparent assets

Endpoints explicitly supporting transparent backgrounds

High-resolution print work

Models supporting 2K or 4K output

Logos and icons

Recraft vector and other SVG-capable models

Large variation batches

Models supporting several outputs per request at manageable cost

Multimodal visual reasoning

Gemini and other models combining analysis with generation

·····

Rapid concept models and final-render models serve different production stages.

A fast model can produce several visual directions while the team is still deciding on composition, scene, mood, or campaign concept.

At that stage, perfect typography, product accuracy, and fine detail may be less important than generating enough alternatives to identify a promising direction.

Once a concept has been approved, a higher-quality model can receive the selected prompt, reference files, negative constraints, and output specifications.

Using the most expensive model at maximum resolution for every early idea increases cost without improving the quality of the creative decision.

Using a fast draft model for the final asset creates the opposite problem, because visible defects, inaccurate products, distorted text, and inconsistent branding may require substantial post-production.

A staged model strategy can therefore reduce spending while preserving quality where it matters.

........

A staged image-production workflow.

Production stage

Suitable model approach

Brief development

Text model converts objectives into visual requirements

Rough exploration

Fast and inexpensive image model generates many concepts

Concept selection

Human reviewer chooses composition and direction

Reference preparation

Approved products, characters, logos, and palettes are collected

Controlled draft

Reference-capable model develops the selected concept

Final render

Higher-quality model and resolution produce the intended asset

Vector production

Specialized model creates scalable SVG where required

Verification

Visual model and human reviewer inspect accuracy

Post-production

Designer completes typography, retouching, colour, and export

·····

Reference images support editing, consistency, and composition.

OpenRouter’s Image API can accept reference images through public URLs or base64-encoded data.

A reference may establish a product shape, character appearance, clothing, architectural style, colour palette, illustration treatment, visual layout, or background environment.

Reference capacity varies substantially among models, with some accepting one image and others supporting more than ten.

A higher reference limit can help when a composition must include several products, multiple character views, a logo, a packaging design, and a separate style board.

The numerical limit does not establish that the model will use every image accurately.

An evaluation must determine whether the model preserves identity, proportions, materials, colours, labels, and relationships after several iterations.

........

Examples of reference-image use in creative workflows.

Workflow

Reference purpose

Product campaign

Preserve packaging, materials, proportions, and logo

Character series

Preserve face, clothing, accessories, and illustration style

Interior visualization

Preserve furniture, materials, and room structure

Architecture

Preserve building form and environmental context

Brand design

Preserve colours, typography, and visual language

Local image edit

Keep unchanged regions while modifying one element

Multi-product layout

Place several approved objects in one composition

Style adaptation

Transfer a visual treatment without copying the exact scene

·····

Reference-image capacity differs widely across the catalogue.

Some OpenAI image endpoints currently accept up to sixteen reference images, while several Gemini and Seedream endpoints allow approximately fourteen.

Certain Riverflow models accept around ten, while several FLUX variants support approximately eight.

Grok Imagine may accept fewer references, while a vector model may use only one.

These numbers can change as providers update their endpoints, so they should be retrieved from the current model record rather than embedded permanently in an application.

The practical decision is not simply whether the endpoint allows enough files, because the total upload size, image resolution, preprocessing, and ability to distinguish the role of each reference also affect performance.

A prompt should explain why each reference exists, such as identifying one as the product, another as the colour palette, and another as the desired environment.

........

Representative differences in reference-image capacity.

Example image family

Approximate current reference limit

GPT Image models

Up to 16

Gemini image models

Up to 14

Seedream 4.5

Up to 14

Riverflow Pro models

Up to 10

Advanced FLUX.2 variants

Up to 8

Grok Imagine Quality

Up to 3

Recraft vector models

Often 1

·····

Image editing should be evaluated according to locality as well as visual quality.

A successful edit changes the requested element while preserving everything that should remain untouched.

A prompt to replace the background should not alter the product label, proportions, face, clothing, pose, or lighting unless those changes are required for visual consistency.

Models vary in their ability to preserve local detail, particularly after repeated edits.

An output may appear attractive while failing the business task because a logo changes, a product acquires an extra component, text is rewritten, or a person’s identity drifts.

Evaluation should therefore compare the edited region and the unchanged regions separately.

A high-quality editing workflow stores the original image, edit prompt, reference files, model, provider, and every intermediate version so that unrequested changes can be identified.

........

Criteria for evaluating image edits.

Evaluation area

Review question

Requested change

Was the intended element modified correctly?

Locality

Did unrelated regions remain stable?

Identity

Did the person, character, or product remain recognizable?

Text

Did existing labels and written content remain accurate?

Materials

Were texture, colour, and surface properties preserved?

Lighting

Does the new element fit the scene naturally?

Composition

Did the edit distort placement or proportions?

Artefacts

Were seams, duplicated objects, or malformed details introduced?

·····

Vector models serve a different purpose from raster image generators.

Raster models produce images made from pixels, which suits photographs, illustrations, social graphics, and visual concepts.

Vector models create paths and shapes that can scale without losing sharpness, making them more appropriate for icons, logos, diagrams, symbols, and certain forms of graphic illustration.

OpenRouter includes Recraft endpoints capable of producing SVG output.

A generated vector may still need professional review because path complexity, typography, spacing, colour use, and trademark risk can make the raw output unsuitable for final identity design.

The availability of SVG reduces the need to reconstruct a raster concept manually, although it does not convert every generated logo into a legally protectable or distinctive brand asset.

........

Raster and vector image workflows compared.

Requirement

Raster output

Vector output

Photorealism

Well suited

Not normally suitable

Painterly illustration

Well suited

Limited

Social-media image

Well suited

Sometimes suitable

Logo concept

Useful for ideation

Better for scalable production

Icon system

Possible but inefficient

Well suited

Infinite scaling

Not available

Available

Pixel-level texture

Strong

Limited

Editing in vector software

Requires tracing

Directly supported

·····

Streaming can make long image renders feel more responsive.

Some image endpoints can return partial renders while the final image is still being created.

An application can display these previews instead of leaving the interface empty for the complete generation period.

Partial images should be labelled as provisional because details, text, objects, and composition may change before completion.

Streaming improves perceived responsiveness rather than guaranteeing shorter total generation time.

A high-resolution or high-quality render may still require a long processing period even when the first preview appears quickly.

The final completed event remains the authoritative output for saving, billing, and review.

·····

Image billing occurs after a completed generation.

OpenRouter states that a completed image is billed in full according to the selected endpoint’s pricing.

A failed or cancelled generation is not billed, while partial streamed previews do not create fractional charges.

This all-or-nothing rule simplifies request accounting, although it does not eliminate the cost of completed outputs that the creative team later rejects.

A technically successful render may still fail because the text is wrong, the product is inaccurate, the subject has visual artefacts, the composition violates the brief, or the style is unsuitable.

Cost analysis should therefore count every completed generation required before one asset is approved.

........

Image-generation costs beyond the advertised endpoint price.

Cost element

Measurement

Completed generations

Every successfully rendered image

Variations

Number of alternatives created before selection

Rejected outputs

Images that cannot be used

Editing rounds

Additional generations required to repair the asset

Prompt-development calls

Text-model cost for brief and prompt creation

Upscaling

Additional model or service cost

Human review

Designer and stakeholder evaluation time

Post-production

Typography, compositing, retouching, and export

Accepted-asset cost

Total workflow cost divided by approved images

·····

Per-image, per-megapixel, and token prices require different calculations.

OpenRouter image endpoints do not use one universal billing unit.

Some charge a fixed amount for each completed image, while others calculate cost according to output megapixels.

Several multimodal models charge input and output tokens, with image generation represented through token-based usage rather than one flat render price.

Higher resolution may create another pricing tier, while reference images, output quality, or provider-specific features can affect the total.

A price displayed as four cents per image can be compared directly with another fixed-image price, although it cannot be compared fairly with a token-priced model without using representative prompts and outputs.

The relevant business metric remains cost per approved asset, which includes every model call and correction required to reach publication quality.

........

Representative image-pricing structures.

Pricing structure

Typical calculation

Per image

Fixed cost multiplied by completed outputs

Per megapixel

Base image cost adjusted for rendered area

Input and output tokens

Prompt and generated-media tokens multiplied by model rates

Resolution tier

Separate price for standard, 2K, or 4K output

Reference surcharge

Additional cost for supplied reference images where applicable

Quality tier

Higher rate for premium rendering settings

·····

OpenRouter can use one model to write the prompt and another to create the image.

The beta image-generation server tool allows a text model to receive a general creative request, determine that an image is required, develop a more detailed visual prompt, and send that prompt to a configured image model.

A user might enter “make a premium coffee advertisement,” after which the orchestrator describes the product position, materials, lighting, camera angle, background, typography area, colour palette, and atmosphere.

The image model then generates the asset from that expanded description.

This two-model structure can help users who know their objective but lack experience writing visual prompts.

It also introduces another source of interpretation, because the text model may change the user’s intention, add an unsuitable style, or omit a requirement.

The original request and enhanced prompt should both be stored so a reviewer can identify whether an error originated in the brief, the prompt-expansion stage, or the image renderer.

........

An orchestrated image-generation workflow.

Stage

Responsible model or user

Initial idea

User provides objective or short description

Brief clarification

Text model identifies missing visual decisions

Prompt enhancement

Text model creates a detailed generation prompt

Image rendering

Selected image model creates the asset

Automated inspection

Vision model checks required elements

Human selection

Reviewer accepts, rejects, or revises the result

Final production

Approved asset receives high-quality render or post-production

·····

Literal and assisted prompt modes should remain separate.

Prompt enhancement is useful when the original request is incomplete, although it may interfere with professional creative direction that is already precise.

A designer who specifies the exact frame, colours, lens, product orientation, text position, negative space, and lighting may not want another model to reinterpret those decisions.

A creative interface can provide an assisted mode for informal prompts and a literal mode that sends the user’s wording to the image endpoint without expansion.

The application may also show the enhanced prompt before rendering, allowing the user to edit or approve it.

This preview improves transparency and prevents an unnecessary generation when the orchestrator misunderstood the objective.

........

Prompt-handling modes for creative applications.

Mode

Behaviour

Assisted

Text model expands a general idea into a detailed visual prompt

Reviewed assistance

Expanded prompt is shown for approval before generation

Literal

Original user prompt is sent without creative rewriting

Template-based

User values are inserted into an approved prompt structure

Professional brief

Detailed creative specification is preserved exactly

·····

Video generation requires a dedicated asynchronous API.

Video rendering can take substantially longer than text or still-image generation, so OpenRouter returns a job rather than holding one request open until the clip is complete.

The application submits the prompt and generation settings to the Video API and immediately receives an identifier and status location.

It then polls periodically or waits for a webhook event until the job becomes completed, failed, cancelled, or expired.

After completion, the application downloads the generated video from an authenticated endpoint.

This architecture requires job storage, status handling, user notifications, timeout rules, and recovery behaviour that are unnecessary for a simple synchronous image request.

........

The OpenRouter video-generation lifecycle.

Job stage

Application behaviour

Submission

Send prompt, model, duration, resolution, and other options

Job creation

Store the returned job identifier

Pending

Inform the user that the request entered the queue

In progress

Poll periodically or wait for a webhook

Completed

Download and store the finished video

Failed

Display the error and decide whether retry is appropriate

Cancelled

Confirm cancellation and update the interface

Expired

Inform the user that the asset is no longer available

·····

Video-generation controls extend beyond a written prompt.

A compatible model may accept duration, resolution, aspect ratio, exact dimensions, frame images, input references, audio settings, seeds, callback URLs, and provider-specific options.

The prompt should describe the subject, action, environment, camera, lighting, timing, and sound rather than providing only a static image description.

A still-image prompt might request a person standing beside a car, while a video prompt should explain how the person moves, how the vehicle behaves, whether the camera tracks or remains fixed, and what should happen during the clip.

The model may also support synchronized dialogue, ambient sound, effects, or music through an audio-generation setting.

These capabilities vary by endpoint and may change the price per second.

........

Common controls available through compatible video endpoints.

Video control

Creative purpose

Prompt

Defines scene, action, camera, timing, lighting, and audio

Duration

Sets the requested clip length

Resolution

Selects 480p, 720p, 1080p, or another supported level

Aspect ratio

Adapts output to landscape, portrait, square, or ultrawide use

Exact dimensions

Requests specific pixel values

First frame

Establishes the exact opening image

Last frame

Establishes the intended final image

Input references

Guides subject identity, objects, environment, or style

Audio generation

Adds synchronized dialogue, ambience, or effects

Seed

Provides limited repeatability where supported

Callback URL

Sends completion events to the application

Provider options

Enables model-specific controls

·····

Frame-based generation and reference-based generation solve different problems.

A frame image defines the exact visual point from which a video begins or ends.

It is useful when the production team has already approved a product shot, character composition, illustration, or storyboard panel and wants the video to animate from that asset.

Reference images guide style, identity, object appearance, or environmental design without becoming a required exact frame.

They are more appropriate when the clip should resemble a visual board while retaining freedom to choose its opening composition.

When both are provided, the API may prioritize frame-image mode because the endpoint must satisfy the stronger first- or last-frame constraint.

........

Frame images and references compared.

Input method

Primary purpose

First-frame image

Fix the exact opening composition

Last-frame image

Fix the intended final composition

First and last frames

Guide a transition between approved endpoints

Style reference

Preserve colour, texture, or illustration treatment

Character reference

Preserve identity, clothing, or proportions

Product reference

Preserve object design and branding

Environment reference

Preserve architecture, setting, and atmosphere

·····

Image-to-video creates a controlled bridge between still design and motion.

One of the most practical workflows begins by generating and approving a still image before paying for video animation.

The still-image stage allows reviewers to confirm character identity, product design, scene composition, lighting, typography area, and overall direction.

The approved asset then becomes the first frame of the video, while the animation prompt focuses on movement, camera behaviour, timing, environmental motion, and sound.

A second approved still may become the final frame where the endpoint supports a controlled transition.

This sequence reduces the number of visual decisions left entirely to the video model, although identity and object consistency may still drift during intermediate frames.

........

A controlled image-to-video pipeline.

Stage

Output

Creative brief

Approved description of scene and motion

Still-image concepts

Several possible opening frames

Human selection

One approved composition

Final still generation

High-quality first frame

Motion prompt

Description of subject, camera, environment, and timing

Video generation

Animated clip beginning from the approved frame

Automated review

Initial checks for continuity and required elements

Human review

Approval of motion, identity, sound, and legal suitability

Post-production

Editing, colour, audio, captions, and export

·····

Reference-to-video supports recurring characters, products, and visual systems.

A campaign may require several clips featuring the same product, fictional character, architectural environment, or visual style.

Reference images can guide the video model toward those recurring elements without forcing every clip to begin from an identical frame.

This mode is useful for product animations, character series, branded stories, interior concepts, architectural presentations, and storyboard-based production.

Consistency should be tested across several outputs, because a model may preserve the general style while changing labels, facial features, proportions, clothing, materials, or object geometry.

The reference set should remain controlled and versioned so different team members do not use conflicting character or product images.

........

Creative uses of reference-to-video generation.

Workflow

Role of the reference set

Product campaign

Preserve product shape, finish, and brand language

Character series

Preserve face, clothing, and illustration treatment

Architectural visualization

Preserve building, materials, and landscape

Interior design

Preserve furniture, room geometry, and palette

Storyboard adaptation

Translate approved panels into moving scenes

Style continuity

Apply the same visual treatment across several clips

Social campaign

Preserve recurring objects and campaign identity

·····

Current video models occupy different speed, quality, and price positions.

OpenRouter’s catalogue includes models from Google, ByteDance, Kuaishou, Alibaba, MiniMax, OpenAI, and other providers.

Google Veo variants emphasize high-fidelity output and synchronized audio, with lighter configurations available for lower-cost iteration.

ByteDance Seedance models focus on reference consistency, camera movement, and controllable video generation, while Fast variants reduce cost and waiting time.

Kling models offer text-to-video and image-to-video options across standard and professional tiers.

OpenAI Sora models provide another production-oriented path, although availability and retirement dates must be monitored.

No single public description establishes the superior model for every workflow, so applications should filter by hard requirements and test representative prompts.

........

Representative video-model positions available through OpenRouter.

Model family

Publicly positioned use

Google Veo

High-fidelity production video and synchronized audio

Veo Lite or Fast variants

Lower-cost ideation and higher-volume generation

ByteDance Seedance

Reference consistency, camera control, and scene generation

Seedance Fast

Faster and less expensive iteration

Kling Standard

General text-to-video and image-to-video

Kling Pro

Higher-quality generation and optional audio

Alibaba Wan

Alternative video generation and model variety

MiniMax Hailuo

Short-form generative video workflows

OpenAI Sora

Production-oriented motion and audio generation

·····

Video price usually scales with duration and output configuration.

Many video endpoints publish a price per generated second.

A ten-second clip priced at five cents per second begins at approximately fifty cents, while a model priced at forty cents per second begins at approximately four dollars for the same duration.

Resolution-specific rates may increase the total for 1080p or another premium output level.

Audio, reference processing, provider options, or model variants may introduce additional pricing differences.

The final campaign cost includes every completed clip generated before approval, not merely the duration of the exported version.

A ten-second approved advertisement may require six rejected clips, several image generations, prompt-development calls, and manual editing.

........

Video-production costs that should be included in budgeting.

Cost component

Measurement

Rendered duration

Model price multiplied by generated seconds

Resolution tier

Additional cost for higher output quality

Draft clips

Early generations used for exploration

Rejected clips

Completed videos that cannot be used

Reference images

Cost of creating or preparing approved frames

Prompt orchestration

Text-model calls used to develop the scene

Audio generation

Additional model or endpoint cost where applicable

Automated evaluation

Vision-model calls used for pre-screening

Human review

Creative, legal, and brand approval time

Post-production

Editing, colour, captions, sound, and export

Accepted-video cost

Total workflow expense divided by approved clips

·····

Video generation can require several minutes rather than one immediate response.

Generation time depends on the model, clip length, resolution, queue, provider load, reference processing, audio, and safety checks.

Applications should communicate job status clearly rather than making the interface appear frozen.

OpenRouter recommends reasonable polling intervals instead of sending constant status requests.

Webhooks are more suitable for production systems because the provider can notify the application when the job reaches a final state.

The webhook should be verified before the application trusts it, preventing a fabricated request from falsely marking a job as complete.

The generated file should be downloaded and stored according to the application’s retention policy before the provider’s temporary asset expires.

·····

Provider-specific video options allow greater control but reduce portability.

Common OpenRouter parameters cover the broad generation workflow, while some providers expose additional controls through passthrough options.

Those settings may include negative prompts, person-generation controls, camera behaviour, safety settings, motion intensity, output configuration, and provider-specific audio options.

An application that depends on one of these settings should inspect the current endpoint record and pin the relevant provider.

Automatic fallback may move the request to another provider that cannot interpret the same option or produces materially different results.

The workflow must decide whether availability or exact creative consistency has priority.

For experimentation, losing one advanced control may be acceptable, while final campaign production may require the tested endpoint with fallbacks disabled.

·····

Provider routing allows creative applications to prioritize price, speed, or consistency.

When several providers serve an eligible model, OpenRouter can restrict the request to a selected list, define a preferred order, exclude endpoints, sort by price, prioritize latency or throughput, and disable fallback.

Price-prioritized routing suits inexpensive ideation when small differences among endpoints are acceptable.

Latency-focused routing suits interactive creative interfaces where users expect a rapid preview.

A pinned provider suits controlled brand production in which every asset must use the endpoint that passed internal testing.

Fallbacks improve availability, although they may change preprocessing, parameter support, visual behaviour, or data handling.

The selected routing policy should therefore appear in the stored generation record.

........

Provider-routing strategies for image and video workflows.

Routing strategy

Appropriate use

Lowest-price endpoint

High-volume exploration and low-risk drafts

Lowest-latency endpoint

Interactive previews and rapid iteration

Highest-throughput endpoint

Large batches or long outputs

Ordered providers

Preferred provider followed by controlled alternatives

Provider restriction

Approved vendors or jurisdictions only

Provider exclusion

Remove endpoints with unsuitable policies or performance

Pinned provider

Final production requiring consistency

No fallback

Prevent untested changes in model behaviour

Fallback enabled

Preserve availability when exact endpoint consistency is secondary

·····

Provider selection also determines where media is processed.

An image or video sent through OpenRouter is transmitted to the provider that ultimately serves the request.

The provider may receive the prompt, reference images, uploaded media, generation settings, and other information needed to complete the task.

OpenRouter’s own retention behaviour and the provider’s retention behaviour remain separate.

OpenRouter states that it normally does not retain prompt and response content unless the account enables private logging or participates in a data-improvement program, while request metadata such as provider, latency, token use, and cost may still be stored.

Providers may apply different retention, abuse-monitoring, training, and geographic-processing policies.

A request requiring stricter privacy should filter endpoints according to their published data policies rather than relying on the model name alone.

........

Media privacy should be evaluated at several layers.

Privacy layer

Review requirement

OpenRouter account settings

Check logging and data-improvement participation

Provider endpoint

Review retention, training, and data-collection policy

Media content

Remove unnecessary personal or confidential information

Reference ownership

Confirm rights to upload and process the files

Processing region

Verify any required geographic restriction

Storage after completion

Define where generated assets will be retained

Access control

Limit who can view prompts, references, and outputs

Publication

Confirm consent, copyright, likeness, and safety requirements

·····

Video generation is incompatible with Zero Data Retention routing.

OpenRouter excludes generated video from Zero Data Retention eligibility because the asynchronous workflow requires temporary storage.

The provider must retain the completed clip long enough for the application to receive the job result and download the media.

A request configured to require ZDR cannot be routed through the Video API.

This limitation may prevent use with confidential prototypes, regulated footage, sensitive likenesses, private customer data, or contracts requiring immediate deletion.

An organization should establish whether temporary provider retention is acceptable before sending reference images or requesting a video.

The absence of ZDR does not automatically mean that a provider trains on the media, although it confirms that some temporary persistence is required.

·····

Image-generation privacy must be checked endpoint by endpoint.

The Image API does not apply one universal public rule stating that every image endpoint is or is not eligible for Zero Data Retention.

The provider record and data-policy controls should therefore determine whether a particular image request satisfies the organization’s requirements.

A model may be available from one provider that collects data and another that does not.

OpenRouter’s routing controls can deny endpoints whose policies permit data collection, although doing so may reduce availability, increase cost, or remove certain parameters.

Sensitive photographs, product prototypes, unpublished campaign material, private documents, and real-person references should receive a more restrictive provider policy than generic fictional ideation.

·····

Real-person references create consent and likeness risks.

OpenRouter provides technical access to image and video models, but the availability of a generation endpoint does not grant the user the right to process or publish another person’s likeness.

A photograph may contain biometric characteristics, location clues, family members, private surroundings, identification documents, confidential products, or copyrighted material.

Generating deceptive, intimate, defamatory, or misleading content involving a real person may violate platform policies, provider rules, privacy law, publicity rights, or other legal obligations.

A model may also alter a person’s face or body in ways that create reputational harm even when the original request appears harmless.

The safest production workflow uses consenting adults, approved brand assets, fictional characters, or properly licensed material.

Legal and human review remain necessary before public distribution.

·····

Generated media may create copyright and trademark uncertainty.

A prompt may request the style of a living artist, a protected fictional character, a recognizable logo, or a product belonging to another company.

The model may generate material that resembles copyrighted or trademarked content without guaranteeing that the output can be used commercially.

OpenRouter routes the generation but does not replace the legal terms of the selected provider or the user’s obligation to hold the necessary rights.

A creative team should maintain a list of prohibited references, protected characters, competitor trademarks, and unlicensed media.

Final assets should be reviewed for accidental logos, copied text, recognizable characters, misleading endorsement, and close similarity to existing campaigns.

A technically successful output remains unusable when the organization cannot establish the right to publish it.

·····

Automated visual evaluation can pre-screen outputs before human review.

A vision-capable model can inspect generated media and answer structured questions about whether required elements appear.

An automated image check may verify the presence of a product, compare visible text with approved wording, identify a distorted logo, or flag an unexpected object.

A video check may summarize the action, inspect the opening and closing frames, describe camera movement, identify visible text, and report whether the intended transition occurred.

This process helps filter large batches before a human reviewer examines the strongest candidates.

Automated review remains probabilistic and may miss subtle identity changes, frame-level artefacts, small text errors, unsafe imagery, or misleading claims.

Human approval remains necessary for public advertising, regulated industries, real-person likenesses, intellectual property, and expensive production.

........

Automated checks that can support creative review.

Review target

Possible automated question

Product presence

Is the required product visible and unobstructed?

Text accuracy

Does the image contain the exact approved wording?

Logo integrity

Is the logo complete and undistorted?

Composition

Is the subject positioned inside the required safe area?

Character consistency

Does the subject resemble the approved reference?

Prohibited content

Does the asset contain an excluded object or theme?

Opening frame

Does the video begin with the approved composition?

Motion

Does the subject perform the requested action?

Final frame

Does the clip end near the approved reference?

Audio

Is dialogue present and synchronized?

·····

Creative evaluation should use a representative set of prompts.

Comparing one impressive demonstration from every model creates a distorted result because each provider can select a prompt suited to its strengths.

A practical test should use several image and video cases drawn from the organization’s real creative work.

Image cases might include a realistic product photograph, a poster containing exact text, a repeated character, a local edit, a transparent asset, a vector icon, a multi-reference composition, and a vertical social design.

Video cases might include a simple camera movement, a product animation, a character action, a transition between approved frames, synchronized dialogue, a vertical social clip, and a scene requiring physical continuity.

Every model should receive the same brief and equivalent parameters where possible.

The evaluator should separate prompt adherence, visual appeal, technical validity, consistency, speed, cost, and required manual correction.

........

Recommended metrics for image and video model evaluation.

Evaluation area

Measurement

Prompt adherence

Required objects, actions, and constraints completed correctly

Visual quality

Detail, lighting, composition, and visible artefacts

Text accuracy

Spelling, wording, hierarchy, and placement

Reference consistency

Preservation of product, person, character, or style

Edit locality

Unrequested regions remain stable

Motion quality

Natural movement, camera behaviour, and continuity

Audio quality

Dialogue, effects, timing, and synchronization

Technical validity

Required format, resolution, ratio, duration, and transparency

Generation latency

Time until the final asset becomes available

Reliability

Successful generations divided by attempts

Raw cost

Total OpenRouter generation spending

Accepted-asset cost

Total spend divided by approved assets

Human correction

Post-production work required before use

Policy suitability

Output can be used under applicable rules and rights

·····

Creative consistency requires storing the complete generation record.

Saving only the final image or video makes it difficult to reproduce an approved asset or understand why another generation differed.

The production record should retain the original brief, enhanced prompt, negative constraints, model slug, endpoint provider, model version, seed, reference files, resolution, aspect ratio, provider options, cost, and generation identifier.

Human approval status should distinguish experimental drafts from material authorized for publication.

The seed should not be treated as a guarantee of exact reproduction because provider infrastructure, model revisions, preprocessing, and nondeterministic operations may still change the result.

The original asset should be stored alongside its metadata rather than assuming that the provider will preserve it indefinitely.

........

Metadata required for reproducible media generation.

Record

Purpose

Original user brief

Preserves the creative objective

Enhanced prompt

Reveals orchestration changes

Negative constraints

Records excluded content

Model slug

Identifies the requested model

Provider endpoint

Identifies the actual serving infrastructure

Model version

Supports later comparison and reproduction

Seed

Assists repeatability where supported

Reference files

Preserves the visual inputs

Resolution and ratio

Records the output dimensions

Provider options

Captures non-standard settings

Generation cost

Supports budget analysis

Generation ID

Connects the asset with OpenRouter activity

Approval status

Separates drafts from publishable media

·····

OpenRouter Presets can encode different creative production profiles.

A Preset can store model choice, provider routing, instructions, tools, and generation parameters under one reusable identifier.

A creative system might maintain a rapid-concept preset that uses a low-cost image model, moderate resolution, and price-prioritized routing.

A final-brand preset might pin one provider, require high quality, preserve approved references, and disable fallbacks.

A video-draft preset could request a short 720p silent clip, while a final-video preset requests 1080p, synchronized audio, webhook delivery, and the tested premium provider.

Central presets allow an organization to change an endpoint or parameter without rewriting every client application.

They also reduce the chance that team members create supposedly comparable assets through different hidden configurations.

........

Example creative presets.

Preset

Intended configuration

concept-fast

Fast image model, moderate resolution, price-prioritized routing

product-edit

Reference-capable model with strict preservation instructions

social-vertical

9:16 format with mobile-safe composition

brand-final

Pinned provider, high quality, approved references, no fallback

vector-export

SVG-capable vector model

video-draft

Fast model, short duration, 720p, no audio

video-final

Premium model, 1080p, audio, webhook completion

character-series

Fixed references, identity instructions, controlled provider

·····

Free and promotional endpoints are appropriate for exploration rather than permanent assumptions.

OpenRouter may expose free variants for certain image or video models.

These endpoints allow developers to test request structures, build interfaces, and compare initial creative behaviour without paying the ordinary generation price.

Free providers may impose lower rate limits, longer queues, changing availability, different retention terms, and temporary promotional conditions.

A production decision should therefore be tested against the paid endpoint expected in deployment.

Latency, reliability, resolution, audio, reference handling, and output quality may differ between free and paid routes even when the model family appears similar.

A free variant can validate the workflow but should not become the sole basis for a commercial budget or service-level commitment.

·····

Model availability and retirement require continuous discovery.

The image and video catalogue changes as providers release new models, promote previews to stable versions, change prices, add endpoints, or retire older slugs.

A model available during development may receive a future removal date before the application reaches production.

Provider availability may also change independently of the model, affecting fallback options, latency, and data policies.

Applications should query OpenRouter’s image and video model endpoints rather than maintaining a permanent manually written list.

The product interface should handle unavailable models, renamed versions, parameter changes, and scheduled retirement without losing the user’s creative project.

Stored presets should also include a replacement strategy for models that disappear.

........

Catalogue information that should be monitored.

Model information

Reason for monitoring

Canonical slug

Required for valid requests

Release date

Helps distinguish current and older models

Retirement date

Provides time to migrate

Endpoint availability

Determines routing and fallback

Pricing

Changes project budgets

Supported parameters

Changes interface controls

Reference limits

Affects consistency workflows

Resolution support

Affects production output

Provider policy

Affects privacy and governance

Preview or stable status

Affects reliability expectations

Content rules

Affects permitted generation

·····

A complete creative workflow should separate experimentation from publication.

The first stage should define the asset, audience, dimensions, visual constraints, references, legal restrictions, and acceptance criteria.

A fast model can generate several low-cost visual directions, after which a human selects the strongest composition.

The approved concept can move to a reference-capable or higher-quality endpoint for detailed generation and local edits.

A final image may become the first frame of a video whose prompt concentrates on motion, camera, timing, and sound.

Automated evaluation can reject obvious failures, while human reviewers confirm brand accuracy, legal suitability, safety, text, identity, and production quality.

The final asset and complete generation record should then be stored before publication.

........

An end-to-end OpenRouter creative workflow.

Stage

Activity

Requirement definition

Set audience, channel, format, rights, and acceptance criteria

Model filtering

Remove endpoints lacking required modalities or parameters

Prompt development

Create or approve a detailed creative brief

Low-cost exploration

Generate several rapid concepts

Human selection

Approve one visual direction

Reference control

Assemble approved products, characters, styles, and frames

Final still generation

Use the tested high-quality endpoint

Video animation

Submit approved frame or references to the Video API

Automated screening

Check visible requirements and obvious defects

Human review

Approve brand, legal, safety, identity, text, motion, and audio

Post-production

Complete editing, colour, typography, sound, and export

Archive

Save output, prompt, references, endpoint, cost, and approval record

·····

Provider choice should be made separately for drafts and final assets.

A price-sorted endpoint may be suitable when the user is exploring ten compositions and expects to discard most of them.

A final advertisement may require a provider that has already passed tests for product consistency, typography, identity preservation, privacy, and legal review.

Using fallbacks during early ideation can reduce failed requests, while disabling them during final generation prevents an untested provider from changing visual behaviour.

An organization may also approve different providers for fictional content and confidential client work.

The routing policy should reflect the consequence of inconsistency rather than applying one default across every creative request.

·····

Accepted-asset cost provides a more useful comparison than model price alone.

A model priced at one cent per image appears inexpensive until the team needs thirty generations and several edits to obtain an acceptable result.

A thirty-cent model may be economically preferable when it creates the approved asset after two attempts.

The same principle applies to video, where one premium clip that follows the storyboard may cost less than several inexpensive clips with inconsistent motion and unusable audio.

The evaluation should divide total workflow cost by the number of assets that passed human approval.

Manual editing and review time should be included when the production decision concerns overall efficiency rather than API spending alone.

........

A complete accepted-asset cost calculation.

Cost category

Included activity

Prompt development

Text-model and human preparation

Draft generation

Every concept image or video

Reference creation

Product, character, style, or frame preparation

Final rendering

High-quality generation request

Editing iterations

Additional model calls for corrections

Automated checks

Vision-model evaluation

Human review

Creative, legal, safety, and brand approval

Post-production

Retouching, typography, editing, audio, and export

Accepted-asset cost

Total cost divided by approved outputs

·····

OpenRouter is most valuable as a multimodal production layer rather than one creative model.

Its common APIs simplify access to image understanding, video understanding, still-image generation, editing, frame-based animation, reference-driven video, and multimodal evaluation.

The platform allows applications to discover current models, inspect provider endpoints, compare pricing units, restrict data policies, use fallbacks, pin providers, and store one activity record across several model developers.

The common interface does not eliminate the differences that determine whether an output is useful, including reference consistency, typography, motion, audio, resolution, privacy, licensing, latency, and manual correction.

Fast endpoints should handle exploration, while approved high-quality endpoints handle final rendering.

Generated stills can become controlled video frames, while vision models can screen outputs before human review.

The operational sequence remains deliberate: the required media output defines the model shortlist, endpoint metadata defines the available controls, provider policy defines where media can be processed, low-cost models create the initial directions, pinned endpoints produce final assets, asynchronous video jobs animate approved references, and human reviewers decide whether the completed material is safe, accurate, legally usable, and ready for publication.

·····

FOLLOW US FOR MORE.

·····

DATA STUDIOS

·····

·····

bottom of page