top of page

Grok Imagine: Generation, Video Creation, Paid Access, API Pricing, Watermarks, Reference Media, and Safety Limits

  • 5 minutes ago
  • 30 min read

Grok Imagine is xAI’s visual-generation system for creating still images, editing existing pictures, combining reference media, animating approved compositions, and producing short videos through both the consumer Grok interface and separately billed developer APIs.

The consumer product emphasizes conversational iteration, allowing users to describe a scene, review the generated result, request changes through follow-up messages, and move from a still image into video without rebuilding the complete creative brief after every generation.

The developer APIs expose a more controlled production environment in which applications select specific image or video models, define resolution, aspect ratio, duration, reference inputs, response format, and output handling while calculating costs according to completed images and generated video seconds.

Free access allows limited experimentation, whereas SuperGrok, X-linked subscription benefits, Grok Business, and enterprise arrangements provide different combinations of higher usage limits, account administration, and product access without converting the service into an unrestricted media generator.

Safety moderation remains active across free accounts, paid subscriptions, business workspaces, and API requests, while watermarks, real-person protections, restrictions involving minors, anti-impersonation rules, intellectual-property obligations, and prohibitions against bypassing safeguards continue to apply regardless of the user’s plan.

The practical distinction therefore runs across four separate environments: casual consumer creation inside Grok, paid consumer use governed by a shared usage allowance, organizational access through Grok Business or Enterprise, and developer production through metered image and video APIs.

·····

Grok Imagine combines image creation, editing, and short-form video within one creative environment.

A user may begin with a written description of a product photograph, fictional character, social-media visual, cinematic scene, poster, interior concept, or illustrated environment, after which Grok generates an initial image that remains available for further conversational direction.

Follow-up instructions may replace the background, adjust the lighting, alter clothing, change the visual style, reposition the subject, modify the atmosphere, or preserve the selected composition while introducing another object or character.

Once the still image reaches an acceptable state, the same creative thread may continue into motion by describing how the subject should move, how the camera should behave, which environmental effects should appear, and whether the clip should include dialogue, ambience, music, or other synchronized sound.

Although the interface presents this sequence as one conversation, the underlying operations remain distinct because image generation, image editing, text-to-video, image-to-video, reference-to-video, video editing, and video extension use different model capabilities, technical constraints, and pricing structures.

........

The principal creative workflows available through Grok Imagine.

Workflow

Starting material

Result

Text-to-image

Written visual description

Newly generated still image

Image editing

Existing image and transformation instruction

Revised image

Multi-reference composition

Several source images and a prompt

New scene combining referenced elements

Visual restyling

Existing image and style direction

Alternative visual treatment

Text-to-video

Written description of scene and movement

Short generated video

Image-to-video

Approved still image and motion prompt

Animated version of the still

Reference-to-video

Subject, object, clothing, or product references

New video guided by the supplied elements

Video editing

Existing short clip and written changes

Modified video

Video extension

Existing clip and continuation instruction

Additional generated footage

·····

Consumer Grok prioritizes conversational direction over exact technical configuration.

The ordinary Grok interface is designed for users who prefer to describe the desired result rather than manage model identifiers, structured parameters, asynchronous jobs, and media-storage procedures.

A prompt may specify subject, environment, mood, lighting, camera angle, materials, colours, text placement, and visual style, while the application chooses many lower-level generation settings automatically.

The same conversation can contain several rounds of revision, allowing the user to preserve the general concept while requesting another expression, background, season, colour palette, composition, or motion sequence.

This interaction model reduces the technical work required for casual creation, although it also reduces predictability when the user needs a fixed model version, exact resolution, repeatable configuration, defined per-generation cost, or consistent reference handling across a large production batch.

Consumer access and API access should therefore be described separately, because the presence of a capability in xAI’s developer documentation does not establish that every Grok subscriber receives the same control through the ordinary app.

........

Consumer Grok Imagine and the developer APIs follow different operating models.

Comparison area

Consumer Grok Imagine

Imagine APIs

Primary audience

Everyday users and individual creators

Developers, production teams, and applications

Interaction model

Conversational prompts and follow-up edits

Structured API requests

Model selection

Simplified or managed by the product

Explicit model identifiers

Image resolution

Managed through the consumer interface, with current marketing up to 2K

Explicit 1K or 2K image selection

Video resolution

Current consumer positioning generally reaches 720p

Supported API workflows reach 1080p

Cost visibility

Shared subscription usage rather than one fixed charge per result

Published per-image and per-second pricing

Reference controls

Simplified through uploads and conversation

Explicit image and video reference parameters

Automation

Limited to the product interface

Integrated through application code

Output handling

Product-managed download and history behaviour

Developer stores temporary URL or base64 output

Moderation

Applied automatically

Applied with moderation information available to the application

·····

Grok Imagine currently separates standard image generation from Quality Mode.

The developer catalogue includes grok-imagine-image, which targets lower-cost generation and editing, and grok-imagine-image-quality, which carries a higher price in exchange for a configuration positioned around greater realism, improved text rendering, and more detailed visual control.

The standard model fits concept development, broad visual exploration, quick social drafts, catalogue variations, and other situations in which the team expects to reject several outputs before selecting a direction.

Quality Mode becomes more appropriate after composition, references, and creative intent have already been approved, because its additional cost is concentrated on the version most likely to enter publication or post-production.

Both models accept written prompts and image inputs, although the presence of reference support does not guarantee exact preservation of a face, product label, logo, material, garment, or other visually sensitive element.

........

Current Grok Imagine image models and output prices.

Model

Intended production position

1K output

2K output

grok-imagine-image

Lower-cost generation, editing, and rapid variation

$0.02 per image

$0.02 per image

grok-imagine-image-quality

Higher-detail generation with stronger realism and text handling

$0.05 per image

$0.07 per image

·····

Reference images create a separate input charge during editing and composition.

A text-only generation pays for the output image, whereas editing and reference-driven requests also charge for each source image supplied to the model.

The current standard model applies a lower input-image rate than Quality Mode, which becomes relevant when a workflow repeatedly submits product photographs, character sheets, campaign references, background images, or earlier generated versions.

These input charges remain small compared with high-resolution video production, although they accumulate across batch generation, repeated edits, and several attempts at preserving the same subject.

A cost estimate should therefore include every reference submitted during the complete creative sequence rather than counting the final output image alone.

........

Current image-input prices.

Model

Price per input image

grok-imagine-image

$0.002

grok-imagine-image-quality

$0.01

·····

The Image API supports formats suited to social media, photography, presentations, banners, and mobile screens.

The available aspect ratios include square, standard landscape, standard portrait, photographic, ultrawide, tall-poster, and smartphone-oriented formats, allowing the same creative concept to be adapted to several publishing surfaces.

A square output suits thumbnails, catalogue products, profile graphics, and conventional social posts, while 16:9 and 9:16 correspond to widescreen and vertical video frames.

Photographic ratios such as 3:2 and 2:3 suit scenes intended to resemble camera output, whereas 2:1, 1:2, and mobile-specific ratios provide additional space for banners, stories, and long-screen layouts.

The API currently supports 1K and 2K image output, with the selected aspect ratio determining the final pixel geometry within the chosen resolution tier.

........

Representative Grok Imagine image formats.

Aspect ratio

Typical publishing use

1:1

Square posts, catalogue items, and thumbnails

16:9

Widescreen scenes, presentation covers, and video frames

9:16

Vertical stories, reels, posters, and mobile video

4:3

Standard presentations and landscape graphics

3:4

Portrait graphics and editorial layouts

3:2

Photography-style landscape compositions

2:3

Photography-style portrait compositions

2:1

Wide banners and header visuals

1:2

Tall posters and narrow editorial graphics

19.5:9

Smartphone-oriented landscape output

9:19.5

Smartphone-oriented portrait output

20:9

Ultrawide mobile composition

9:20

Extra-tall mobile composition

Auto

Aspect ratio selected by the model

·····

Batch image generation produces several alternatives from one creative brief.

The API can return as many as ten image variations within one request, allowing an application to create a broad concept set before the user begins detailed editing.

This approach suits campaign ideation, catalogue imagery, character development, background selection, style comparisons, and automated workflows in which another model or human reviewer scores the alternatives.

Every completed image is billed, so a ten-image batch represents ten generated outputs rather than one output accompanied by nine free variations.

The number of requested alternatives should reflect the expected rejection rate and the cost of reviewing them, because a larger batch may create more selection work without producing a composition that satisfies the brief.

·····

Cost per approved image reflects failed concepts, revisions, and human correction.

The listed output price describes the model charge for one completed image, whereas the production cost includes all earlier concepts, reference inputs, quality rerenders, local edits, rejected results, automated checks, and manual post-production.

A campaign visual may begin with ten standard-model drafts, continue through several reference-based edits, and end with two 2K Quality Mode renders before one version receives approval.

An image model with a higher output price may still produce a lower cost per approved asset when it preserves text, products, and composition accurately enough to avoid repeated correction.

The appropriate comparison therefore divides the full creative expenditure by the number of images that pass final review rather than ranking models according to one published generation price.

........

Costs that contribute to one approved Grok Imagine image.

Cost category

Included activity

Concept generation

Standard-model variations created during exploration

Reference processing

Product, character, style, and environmental inputs

Quality rerendering

Higher-resolution or Quality Mode output

Editing rounds

Additional generations used to correct the image

Rejected results

Completed images that cannot be published

Automated inspection

Vision-model checks for text, objects, or branding

Human approval

Creative, legal, safety, and marketing review

Post-production

Typography, retouching, compositing, and export

Approved-asset cost

Total workflow expenditure divided by accepted images

·····

Natural-language editing recreates portions of the image rather than applying deterministic pixel changes.

A user may request a new background, different clothing, altered lighting, another artistic treatment, changed weather, revised composition, or an additional object while expecting all unrelated elements to remain stable.

The generation model may reconstruct a wider portion of the scene than the written instruction appears to require, which can alter faces, hands, labels, logos, proportions, textures, lighting, or small environmental details.

A result may therefore look visually coherent while failing the operational objective because the requested background changed correctly but the product packaging no longer matches the approved reference.

Review should compare the edited output with the source image and classify both the intended transformation and every unintended difference that affects identity, branding, text, or factual content.

........

Image-editing review should separate requested and unrequested changes.

Review area

Technical question

Requested transformation

Did the model perform the specified edit?

Locality

Did regions outside the intended edit remain stable?

Subject identity

Does the person or fictional character remain recognizable?

Product preservation

Are dimensions, materials, packaging, and logo unchanged?

Written content

Are labels, captions, and signs still accurate?

Composition

Did object placement or scale change unexpectedly?

Lighting integration

Does the edited element fit the scene naturally?

Generation artefacts

Were malformed objects, edges, hands, or faces introduced?

·····

Multi-image editing accepts up to three source references.

The current editing workflow allows as many as three images to contribute to one generated composition, which supports combinations such as product, environment, and campaign style or character face, clothing, and scene reference.

A fashion concept may combine a person, garment, and setting, while a product advertisement may use front packaging, side packaging, and a lifestyle background.

Three inputs remain restrictive for complex productions containing several products, multiple character views, accessories, typography references, logos, and detailed style boards.

Such projects may require a staged workflow in which one generation produces an approved composite reference that becomes the source for the next edit.

........

Examples of three-reference composition workflows.

First reference

Second reference

Third reference

Intended result

Product photograph

Environmental background

Brand-style reference

Campaign visual

Character portrait

Clothing reference

Scene reference

Character illustration

Person

Object

Artistic treatment

Stylized composition

Existing interior

Furniture

Colour palette

Redesigned room

Package front

Package side

Lifestyle setting

Product advertisement

Illustration

Texture

Lighting reference

Restyled artwork

·····

Image-to-video preserves an approved opening composition before motion begins.

Creating the still image first allows the user to verify subject appearance, product design, environment, framing, lighting, colour palette, and visual style before paying for animation.

The approved image then becomes the opening frame, while the video prompt concentrates on movement, camera direction, environmental effects, physical interaction, pacing, dialogue, and sound.

This sequence reduces the number of visual decisions generated simultaneously, although it does not guarantee that every facial feature, product detail, garment, or background element will remain unchanged during motion.

Review should inspect multiple frames across the complete clip rather than relying on the preview thumbnail or first frame.

........

A controlled image-to-video sequence.

Production stage

Creative decision

Still-image brief

Define subject, setting, style, and composition

Concept generation

Produce several possible opening images

Human selection

Approve one composition

Quality rendering

Create the final detailed still

Motion specification

Define subject movement, camera, timing, and sound

Video generation

Animate the approved image

Frame review

Check continuity, identity, products, and background

Audio review

Check speech, effects, ambience, and synchronization

Post-production

Add captions, transitions, colour correction, and final sound

·····

Grok Imagine Video supports several generation and transformation modes.

Text-to-video asks the model to invent the scene, characters, objects, framing, motion, and sound from written instructions, leaving the largest number of creative decisions to generation.

Image-to-video uses one approved visual as the opening state, providing greater control over initial composition and subject appearance.

Reference-to-video guides the new clip with people, products, garments, objects, or visual styles without requiring those references to become the exact first frame.

Video editing modifies an existing short clip through natural-language instructions, while video extension continues the motion and narrative beyond the original ending.

........

Grok Imagine video modes follow different creative constraints.

Video mode

Starting material

Production purpose

Text-to-video

Written scene description

Generate an entire scene from language

Image-to-video

One approved still image

Preserve the opening composition

Reference-to-video

Subject, product, clothing, or style references

Guide a newly generated scene

Video editing

Existing short video

Modify selected visual or narrative elements

Video extension

Existing ending

Continue the generated sequence

·····

Video 1.5 currently supports clips lasting from one to fifteen seconds.

The duration range suits social-media segments, product moments, cinematic transitions, animated illustrations, short demonstrations, and brief advertising concepts.

A longer sequence requires several separate generations that are later assembled in conventional editing software.

Continuity becomes more difficult when each clip is produced independently, because facial details, clothing, object position, camera language, lighting, and environment may shift between requests.

The final frame of one approved segment may serve as the source image for the next segment, although the production team must still correct discontinuities during editing.

·····

Consumer video resolution and API video resolution follow different limits.

The consumer Grok product currently presents text-to-video creation generally at up to 720p, while the grok-imagine-video-1.5 API supports 480p, 720p, and 1080p for compatible text-to-video and image-to-video requests.

Reference-to-video currently remains capped at 720p, and video editing also returns output at no more than 720p even when the original clip has a higher resolution.

A paid consumer subscription should therefore not be described as including every 1080p developer capability, because subscription access and metered API access remain separate products.

........

Current documented video-resolution limits.

Workflow

Maximum documented resolution

Consumer Grok text-to-video

Generally advertised up to 720p

API text-to-video with Video 1.5

1080p

API image-to-video with Video 1.5

1080p

API reference-to-video

720p

API video editing

720p

Legacy video-model output

Limited to the resolutions published for that model

·····

Video 1.5 generates visual motion and sound within the same production pass.

xAI positions Video 1.5 around more coherent subject movement, more believable physical weight and momentum, clearer speech, and tighter synchronization between visual action and audio.

A prompt may specify spoken dialogue, ambient sound, music, environmental noise, or effects alongside the scene and camera behaviour.

Generated audio still requires separate inspection because speech may contain incorrect words, timing may drift, accents may not suit the character, and environmental sound may conflict with the visual setting.

The complete asset should therefore be approved as an audiovisual sequence rather than as a silent image stream accompanied by automatically trusted sound.

........

Video 1.5 requires combined visual and audio review.

Review area

Possible failure

Subject movement

Unnatural or inconsistent motion

Object physics

Incorrect weight, collision, or momentum

Camera behaviour

Unwanted shaking, drifting, or impossible movement

Identity continuity

Face, hair, clothing, or body changes

Product continuity

Packaging, logo, or geometry drifts

Speech accuracy

Spoken words differ from the prompt

Lip synchronization

Mouth movement fails to match speech

Environmental sound

Audio does not correspond with the scene

Music

Mood or timing conflicts with the brief

Scene continuity

Objects appear, disappear, or transform unexpectedly

·····

Reference-to-video guides recurring people, products, clothing, and other visual elements.

A brand may supply a product image and request a completely new environment, while a fictional-character project may provide approved portraits and ask for another scene that preserves face, costume, and visual treatment.

Fashion workflows may use garment and subject references, while automotive, architectural, and interior projects may depend on the preservation of recognizable objects or structures.

The model attempts to retain the referenced characteristics, although exact logos, measurements, facial details, text, and proportions may change during generation.

Production review should inspect representative frames throughout the clip, especially when the referenced subject carries legal, commercial, or identity significance.

........

Reference-to-video workflows require subject-specific verification.

Workflow

Referenced properties requiring review

Product animation

Shape, finish, label, logo, and materials

Character sequence

Face, clothing, accessories, and illustration style

Virtual styling

Garment appearance, subject identity, and fit

Automotive scene

Vehicle model, body details, and branding

Architectural visualization

Building form, materials, and environment

Branded social clip

Campaign colours, objects, and product identity

Fictional-series production

Character and location continuity across clips

·····

Reference audio remains restricted and does not provide unrestricted voice cloning.

Current documentation describes the use of preset xAI voices in certain reference-to-video workflows, with as many as three supported voice identifiers included in one request.

The feature does not allow an ordinary user to upload any personal voice recording and assume that Grok will reproduce it as a cloned voice.

Reference-audio access is currently restricted to trusted partners in the United States, so consumer users and most API developers should not treat it as a standard part of Video 1.5.

Generated clips still include audio capabilities without that restricted reference-audio workflow.

·····

Video editing preserves the original duration and aspect ratio within stricter limits.

The editing endpoint accepts an existing short clip and a written description of the requested change, while the generated result retains the source duration rather than accepting a new duration parameter.

The current source-duration limit is approximately 8.7 seconds, which excludes longer clips from one-pass editing.

The output follows the original aspect ratio and is capped at 720p, meaning that a 1080p source may return at a lower resolution after editing.

These constraints make the endpoint suitable for focused modifications to short clips rather than complete editing of a long production.

........

Current Grok Imagine video-editing constraints.

Editing property

Current treatment

Maximum source duration

Approximately 8.7 seconds

Custom output duration

Not supported

Output duration

Matches the source clip

Custom aspect ratio

Not supported

Output aspect ratio

Matches the source clip

Maximum output resolution

720p

Natural-language editing

Supported

Audio processing

Included where the workflow supports it

·····

Video API requests complete asynchronously rather than returning the finished clip immediately.

The application submits the model, prompt, duration, resolution, aspect ratio, and references, after which xAI returns a request identifier for later status checks.

Processing may require several minutes according to prompt complexity, clip length, resolution, input media, system demand, and the requested video operation.

The application must display pending, completed, failed, or expired states instead of presenting the request as an ordinary synchronous response.

Completed video URLs remain temporary, so the application should download or transfer the asset before the hosted result expires.

The moderation result should also be inspected before the file enters user-visible storage or publication workflows.

........

The Grok Imagine video job lifecycle.

Job stage

Required application behaviour

Submission

Send prompt, model, duration, ratio, resolution, and references

Identifier return

Store the request ID

Processing

Poll at a reasonable interval

Completion

Retrieve moderation information and output URL

Download

Save the video before the temporary URL expires

Review

Inspect image quality, motion, sound, safety, and rights

Archival

Store the final media and complete generation metadata

Failure or expiry

Present an error without attempting safety circumvention

·····

API video pricing increases with duration and resolution.

The current Video 1.5 model charges for every generated second, with separate rates for 480p, 720p, and 1080p output.

Input reference images add a smaller per-image charge.

The older grok-imagine-video model carries lower per-second rates but does not list the same 1080p output option.

Every successfully completed clip incurs the published charge even when the team later rejects it because of visual drift, incorrect dialogue, unsuitable motion, or other creative defects.

........

Current Grok Imagine video API pricing.

Model

Input-media charge

480p output

720p output

1080p output

grok-imagine-video-1.5

$0.01 per input image

$0.08 per second

$0.14 per second

$0.25 per second

grok-imagine-video

$0.002 per input image or $0.01 per second of video input

$0.05 per second

$0.07 per second

Not currently listed

·····

A fifteen-second 1080p Video 1.5 generation costs $3.75 before reference charges.

At the current published rate, five seconds of 480p output costs forty cents, while fifteen seconds costs one dollar and twenty cents.

The same durations at 720p cost seventy cents and two dollars and ten cents, whereas 1080p output costs one dollar and twenty-five cents for five seconds and three dollars and seventy-five cents for fifteen seconds.

These figures describe one completed output and exclude reference-image fees, still-image development, rejected attempts, automated inspection, and manual post-production.

........

Representative Video 1.5 output costs.

Duration

480p

720p

1080p

5 seconds

$0.40

$0.70

$1.25

10 seconds

$0.80

$1.40

$2.50

15 seconds

$1.20

$2.10

$3.75

·····

The cost of an approved video includes every unsuccessful creative attempt.

A fifteen-second 1080p output may carry a direct model cost of $3.75, although the complete production may include several still-image concepts, one or more Quality Mode renders, multiple video attempts, reference processing, automated checks, and manual editing.

One clip may be rejected because a face changes, another because a product label becomes unreadable, and another because the generated dialogue contains the wrong wording.

Post-production may add captions, transitions, sound correction, colour grading, and brand elements after the generated version has been selected.

The relevant comparison divides total expenditure by the number of clips approved for publication rather than treating the final duration as the complete budget.

........

Costs contributing to one approved Grok Imagine video.

Cost category

Included activity

Creative planning

Human or model-assisted prompt development

Still-image concepts

Draft opening frames and reference development

Quality still

Approved high-resolution source image

Video generations

Every completed clip created during iteration

Reference inputs

Product, character, clothing, or style images

Automated checks

Vision-model screening for visible errors

Human approval

Creative, legal, safety, and brand review

Post-production

Editing, colour, sound, captions, and export

Approved-video cost

Total expenditure divided by accepted clips

·····

Free access provides experimentation without a guaranteed production allowance.

The consumer Grok product allows free users to try image and video generation within the limits applied to their accounts.

Availability may vary according to system demand, geographic rollout, platform, temporary restrictions, and the particular visual feature being used.

A free account should therefore be treated as an experimentation route rather than as a production service with a permanent daily number of images or videos.

The practical capacity required for frequent creative work normally comes from a paid Grok plan or separately funded developer API account.

·····

SuperGrok increases the shared weekly allowance rather than removing usage limits.

Paid consumer activity is currently measured through one weekly usage pool covering Chat, Imagine, Voice, Build, and other eligible Grok functions.

A basic conversation consumes comparatively little compute, while a high-resolution video may use a substantial part of the same allowance.

Two subscribers on the same plan may therefore generate different numbers of images or videos according to duration, resolution, quality settings, and activity in other Grok products.

The Usage area reports the percentage consumed, the product breakdown, the next reset date, and any Extra Usage Credits available to the account.

........

Factors affecting a paid user’s practical Imagine capacity.

Usage factor

Technical consequence

Subscription tier

Determines the size of the included weekly pool

Image quality

Higher-quality generation may consume more allowance

Video duration

Longer clips require greater compute allocation

Video resolution

Higher resolution consumes more allowance

Other Grok products

Chat, Voice, and Build draw from the same pool

Temporary capacity limits

Product demand may alter availability

Purchased Extra Usage Credits

Extend activity after the included pool is exhausted

Weekly reset

Restores the included allowance on the account’s schedule

·····

Paid plans do not include one permanently published number of image or video generations.

A fixed statement such as a guaranteed monthly video count would conflict with the current shared-usage model.

A subscriber who creates short 480p drafts and uses little of the wider Grok product may complete more generations than another subscriber who creates several 720p videos while also using Build and Voice extensively.

Temporary limits, regional availability, and changing product rules may also alter practical capacity.

The live Usage screen remains the account-specific source for remaining allowance rather than a universal generation count published for every subscriber.

·····

Consumer 720p video may fall back to 480p after the higher-resolution allowance is exhausted.

The current consumer documentation describes a fallback process in which a requested 720p video may be generated at 480p after the relevant account limit has been reached.

A successful request can therefore return a technically different result from earlier generations during the same usage period.

Creators producing a consistent series should inspect the resolution of every downloaded clip rather than assuming that the requested quality remained available.

The API provides more explicit control because the developer selects the resolution and pays the corresponding per-second rate.

·····

Extra Usage Credits extend paid consumer access without becoming API credits.

Eligible paid users may purchase Extra Usage Credits after consuming the included weekly allowance, and automatic top-ups may also be available through account settings.

These credits remain separate from the weekly reset and currently expire after one year unless the purchase terms specify another period.

Users should confirm the email address, Apple account, Google account, or linked X identity associated with the purchase because credits and subscription benefits belong to the signed-in account.

Consumer usage credits do not fund developer API calls, while API credits do not convert the consumer account into SuperGrok.

·····

X-linked Grok benefits and standalone SuperGrok subscriptions require account alignment.

Grok may recognize benefits connected to an eligible X subscription when the X identity is linked correctly to the xAI account.

X Premium and Premium+ remain administered through X, whereas standalone SuperGrok billing follows xAI’s consumer subscription system.

Benefits apply to the connected account rather than automatically extending across every identity owned by the subscriber.

Users who subscribe through several platforms should verify the active account before purchasing another plan to resolve apparently missing access.

·····

Grok Business adds organizational administration without replacing API billing.

Grok Business currently includes Imagine within a managed team workspace, together with higher quotas, role-based access, consolidated billing, team administration, and a commitment that business data is not used for model training.

The public price is currently thirty dollars per user per month, while Enterprise arrangements follow negotiated contracts.

A Business seat provides product access for team members, although separately metered developer API generation remains governed by API credentials, credits, and contractual terms unless the organization’s agreement states otherwise.

The distinction matters when a company wants both collaborative creative use inside Grok and automated image or video generation inside its own software.

........

Grok Imagine access routes follow separate billing structures.

Access route

Primary purpose

Charging model

Free Grok

Limited individual experimentation

No subscription within current limits

SuperGrok

Frequent personal use across Grok products

Subscription with shared weekly usage

X-linked Grok benefits

Grok access connected to an X subscription

Managed through X billing

Grok Business

Team workspace and organizational controls

Per-user subscription

Grok Enterprise

Customized organizational deployment

Contract pricing

Imagine APIs

Application integration and production automation

Per image and per generated video second

·····

API credits and consumer subscriptions remain financially separate.

A SuperGrok subscription does not provide an unlimited Imagine API balance, while purchasing developer credits does not upgrade the consumer account.

Applications using grok-imagine-image, grok-imagine-image-quality, or grok-imagine-video-1.5 should calculate costs from the published API rates.

Consumer users should monitor the weekly shared allowance rather than applying per-image or per-second API prices to their subscription activity.

xAI describes API credits as non-refundable, so developers should begin with a limited balance while testing moderation behaviour, output quality, latency, storage handling, and accepted-asset cost.

·····

NSFW preferences alter some adult-content availability without disabling moderation.

Consumer settings may expand access to certain mature or adult-oriented material under eligible account and regional conditions.

The setting does not remove model safeguards, account enforcement, legal restrictions, or the xAI Acceptable Use Policy.

Prohibited categories remain blocked across free accounts, SuperGrok, higher consumer tiers, business workspaces, and API requests.

xAI does not publish a complete dictionary of blocked words or classifier thresholds, because detailed bypass guidance would facilitate circumvention and the moderation systems change over time.

........

Moderation remains active across every current access category.

Access or setting

Safety moderation

Free consumer account

Active

SuperGrok

Active

Higher paid consumer tier

Active

X Premium-linked access

Active

Grok Business

Active

Grok Enterprise

Active

Imagine API

Active

NSFW preference enabled

Active

Extra credits purchased

Active

High-resolution generation

Active

·····

Sexual content involving minors remains prohibited under every setting and visual style.

The restriction covers real people, fictional characters, illustrations, animation, photorealistic generations, edited images, and synthetic videos.

Describing an apparently young character as technically adult does not create a permitted request when the visual or contextual presentation remains minor-like.

Repeated attempts may trigger account restrictions or suspension, while material that appears to contain child sexual abuse may be reported to the appropriate authorities.

The same prohibition continues when the user supplies the source image, enables NSFW settings, purchases a subscription, or states that the output will remain private.

·····

Real-person nudification and non-consensual intimate imagery are prohibited.

The xAI Acceptable Use Policy prohibits requests that remove clothing from a real person, insert a real person into an intimate or pornographic setting, or generate sexual material without the depicted person’s consent.

The rule applies to celebrities, public figures, colleagues, former partners, private individuals, and anyone whose photograph can be found online or uploaded by the user.

Public availability of an image does not create permission for sexual transformation, while private intended use does not replace the missing consent of the depicted person.

Reference-image editing therefore requires both the technical right to use the image and an appropriate purpose for the requested transformation.

........

Prohibited real-person intimate manipulations.

Requested transformation

Reason for prohibition

Remove clothing from a real photograph

Nudification

Place a celebrity in an explicit scene

Non-consensual intimate imagery

Sexualize a colleague’s profile photograph

Privacy and consent violation

Generate an intimate deepfake of a former partner

Non-consensual abuse

Convert a public photograph into pornography

Likeness and consent violation

Create explicit media from a private person’s reference

Exploitation and privacy violation

·····

Deceptive impersonation, forged evidence, and false endorsements remain restricted.

The Acceptable Use Policy prohibits material misrepresentation involving identity, credentials, signatures, evidence, documents, records, and endorsements.

A generated photograph of a public figure supporting a product may mislead viewers even when the scene is technically fictional.

Fabricated identification cards, false certificates, forged signatures, fake crime scenes, and misleading news photographs create risks involving fraud, defamation, false light, and public deception.

Satirical or fictional work should include enough context to prevent the audience from treating the generated scene as authentic evidence.

........

Generated-media scenarios carrying impersonation or deception risks.

Generated output

Principal concern

False celebrity endorsement

Misrepresentation and publicity rights

Fabricated news image

False evidence

Fake identity document

Credential fraud

Forged signature

Impersonation and fraud

False criminal depiction

Defamation and false light

Fake professional qualification

Fabricated record

Misleading political-event image

Public deception

Synthetic customer testimonial

False endorsement

·····

Copyright, trademarks, privacy, and publicity rights remain the user’s responsibility.

A successful generation confirms that the model returned an output, although it does not establish that the user owns every right required for publication, advertising, merchandise, resale, or other commercial use.

Prompts involving copyrighted characters, competitor packaging, recognizable logos, real people, living artists, or protected photographs may create material that requires licensing or may remain unsuitable for public distribution.

Uploaded references may belong to photographers, designers, employers, publishers, clients, or other rights holders whose permission has not been granted.

Commercial teams should review the source material, generated result, provider terms, intended audience, and distribution channel before approving publication.

........

Creative scenarios requiring additional rights review.

Creative request

Potential legal or contractual issue

Recognizable fictional character

Copyright and trademark

Competitor branding or packaging

Trademark and unfair competition

Living artist’s identifiable style

Attribution and copyright concerns

Real person in an advertisement

Publicity and consent

Uploaded professional photograph

Photographer’s copyright

Confidential product prototype

Trade secrets and contractual confidentiality

Celebrity voice or appearance

Likeness and endorsement

Generated merchandise design

Commercial-use and derivative-work rights

·····

Consumer outputs retain the Grok watermark regardless of paid status.

Images and videos created through the consumer product include a Grok watermark identifying them as AI-generated.

The ordinary consumer interface does not provide a control for removing it, while SuperGrok does not publicly promise an unwatermarked mode.

The Acceptable Use Policy separately prohibits hiding, stripping, altering, or circumventing embedded watermarks and provenance signals.

Organizations requiring another output treatment should evaluate the relevant API or enterprise terms rather than cropping or editing the consumer watermark away.

........

Current watermark expectations across Grok Imagine access routes.

Generation context

Watermark treatment

Free consumer generation

Remains present

SuperGrok generation

Remains present

Higher consumer subscription

No public promise of removal

NSFW preference enabled

Remains present

Downloaded consumer image

Remains present

Edited consumer output

Provenance rules continue

API or enterprise output

Governed by the applicable technical and contractual terms

Manual removal attempt

Prohibited

·····

X publication rules and Grok generation rules operate independently.

X may permit certain consensual adult media when it is labelled correctly, restricted appropriately, and placed outside prohibited profile surfaces.

That publication policy does not require Grok Imagine to generate the same media, because xAI moderation determines whether creation is permitted before X rules determine whether an existing file may be posted.

A user may therefore possess media that X allows under a warning label while Grok refuses to create it.

Paid account status does not merge these two policy systems or provide a route around generation safeguards.

........

Generation and publication follow different policy layers.

Question

Governing policy

Will Grok generate the image?

xAI moderation and Acceptable Use Policy

Will Grok generate the video?

xAI moderation and model safeguards

Can the completed file be posted on X?

X Rules

Does the post require a sensitive-media label?

X content-label rules

Can the media appear in a profile image or header?

X placement restrictions

Does a paid plan bypass Grok moderation?

No

Does X posting permission guarantee Grok generation?

No

·····

Attempts to circumvent moderation create a separate policy violation.

Changing spelling, encoding instructions, translating a blocked prompt, splitting one prohibited transformation into several smaller edits, or repeatedly resubmitting the same request may be treated as deliberate safety evasion.

The Acceptable Use Policy prohibits jailbreaking, adversarial prompting, limit circumvention, and attempts to bypass protective controls.

Repeated refusals should therefore be treated as a boundary rather than as an invitation to search for another formulation that produces the same prohibited result.

Account restrictions may follow patterns of repeated circumvention even when each individual request is phrased differently.

·····

API developers receive moderation information and retain responsibility for the surrounding product.

Image and video responses include moderation-related information that the application should inspect before displaying, storing, or publishing the generated media.

A passed result may continue into creative and legal review, whereas a filtered result should remain hidden from the user.

Repeated prohibited requests may require account-level restrictions, manual investigation, or additional safeguards within the developer’s application.

Upstream moderation does not replace the need for local terms, user reporting, abuse detection, rate controls, audit records, and human escalation procedures.

........

Application responses should follow the moderation outcome.

Moderation outcome

Appropriate application action

Passed

Continue to quality, rights, and publication review

Filtered

Do not display the generated media

Request rejected

Present an appropriate error without suggesting circumvention

Repeated prohibited requests

Apply account-level controls

Ambiguous visual result

Route to human review

Reported content

Preserve relevant evidence and investigate

Safety-system uncertainty

Prevent automatic publication until reviewed

·····

Grok is not intended for children under thirteen, while teenagers require guardian permission.

Users between thirteen and seventeen require consent from a parent or guardian, and xAI warns that Grok may produce coarse language, sexual situations, violence, or other material unsuitable for some younger audiences.

The general age threshold does not provide unrestricted access to mature features, because account controls, regional law, product safeguards, and X age restrictions continue to apply.

Parents and guardians should review data controls and supervise younger users when visual generation forms part of the account’s activity.

The prohibition against sexual content involving minors remains absolute and does not depend on the user’s age.

·····

Imagine API media receives content review without being used for model training.

xAI’s current developer documentation states that media processed through the Imagine API is reviewed for policy compliance and is not used to train models.

This commitment applies to the developer product and should not automatically be extended to every consumer data configuration.

Consumer users have separate Grok data controls that determine how account content may contribute to product improvement.

Enterprise arrangements may add regional processing, single sign-on, role-based access, audit logging, and other contractual controls for eligible organizations.

·····

Zero Data Retention changes how generated media must be returned and stored.

Under the documented ZDR workflow, image output must be returned as base64 data rather than through an xAI-hosted temporary URL.

Video output must be transferred to a customer-provided upload destination, while xAI-hosted stored output and persistent file identifiers are unavailable.

The Console video playground cannot be used for ZDR video creation, and certain agentic image workflows remain outside the supported ZDR path.

Content moderation continues even when storage is minimized, because ZDR changes retention architecture rather than safety policy.

........

Imagine API behaviour under Zero Data Retention.

Feature

ZDR requirement

Hosted image URL

Unavailable

Base64 image output

Required

xAI-hosted video result

Unavailable

Customer-provided video upload URL

Required

Persistent xAI file identifier

Unavailable

Console video playground

Unavailable

Certain agentic image workflows

Not supported in the documented ZDR route

Safety moderation

Continues

·····

Temporary output URLs require deliberate download and archival procedures.

Standard image and video API responses may provide URLs that expire rather than permanent media storage.

Applications should transfer the completed asset into approved storage as soon as the generation and moderation checks finish.

The production record should preserve the original prompt, follow-up edits, model identifier, resolution, ratio, duration, reference inputs, output cost, moderation status, generation identifier, downloaded file, and human approval decision.

Exact regeneration may remain impossible even when the same prompt and settings are reused, making the original downloaded asset part of the required audit record.

........

Metadata worth preserving with every completed asset.

Stored record

Operational purpose

Original prompt

Preserves the first creative instruction

Follow-up edit prompts

Records how the asset changed

Model identifier

Distinguishes standard, Quality Mode, and video generation

Resolution

Records technical output quality

Aspect ratio

Records composition format

Duration

Records video length

Reference inputs

Preserves the source material

Output cost

Supports budget analysis

Moderation status

Documents safety handling

Generation identifier

Supports technical investigation

Downloaded media file

Prevents loss after URL expiry

Approval decision

Separates experimental drafts from publishable assets

·····

Reference photographs should be minimized, licensed, and supplied with consent.

An uploaded photograph may contain a face, home interior, workplace, identification document, family member, private object, location clue, or confidential product design that has no relevance to the requested transformation.

Removing unnecessary information reduces privacy exposure and limits the amount of sensitive material processed during generation.

Real-person references should be used only with appropriate consent and rights, especially when the output changes clothing, setting, age appearance, emotional context, or commercial association.

Organizations should also restrict access to the original references and completed media rather than placing all creative material in a broadly shared storage location.

·····

A professional production sequence should reserve expensive generation for approved directions.

The standard image model can produce several low-cost concepts from the initial brief, allowing the team to compare composition, subject, colour palette, and visual style before committing to a higher-quality render.

Approved products, characters, environments, and other references should then be assembled into a controlled set that remains consistent across later requests.

A 2K Quality Mode image can become the source for Video 1.5, while early animation tests may use 480p or 720p before the team pays for the final 1080p output.

Automated checks may detect obvious text, logo, object, or subject errors, although human reviewers remain responsible for legal rights, identity preservation, brand accuracy, safety, motion quality, sound, and publication context.

........

A cost-controlled Grok Imagine production workflow.

Production stage

Recommended operation

Brief definition

Establish audience, channel, format, rights, and safety constraints

Concept generation

Use the standard image model for several visual directions

Direction selection

Approve composition, subject, and style

Reference control

Assemble licensed product, character, and environmental inputs

Final still

Produce a 2K Quality Mode image

Motion test

Generate a short 480p or 720p video

Revision

Correct movement, camera, identity, product details, and sound

Final video

Generate the approved duration and resolution

Review

Inspect text, branding, continuity, safety, and legal rights

Archive

Store prompts, references, metadata, moderation result, and media

Publication

Preserve watermarks and apply any additional disclosure

·····

Representative evaluation prompts expose differences that marketing examples conceal.

A credible model comparison should use several tasks resembling the organization’s real creative work rather than one demonstration selected because it suits the model.

Image cases may include a realistic product scene, a poster containing exact text, a fictional character appearing in several environments, a local edit, a three-reference composition, and a vertical social visual.

Video cases may include a defined camera movement, product animation, dialogue scene, character action, reference-driven clip, and motion sequence requiring consistent physical interaction.

Each result should be scored for prompt adherence, visual detail, text accuracy, identity preservation, edit locality, motion, physical coherence, sound, technical validity, latency, moderation behaviour, and human correction.

........

Recommended Grok Imagine evaluation metrics.

Evaluation area

Measurement

Prompt adherence

Required subjects, actions, and constraints appear correctly

Visual quality

Detail, lighting, composition, and artefacts

Text accuracy

Spelling, wording, hierarchy, and placement

Reference consistency

Preservation of person, product, object, or style

Edit locality

Unrequested regions remain stable

Motion quality

Subject and camera movement remain natural

Physical coherence

Weight, momentum, interaction, and continuity

Audio quality

Dialogue, sound, timing, and synchronization

Technical validity

Correct duration, ratio, resolution, and file format

Generation latency

Time from submission to completed output

Moderation behaviour

Appropriate filtering of prohibited requests

Raw model cost

Output and reference charges

Approved-asset cost

Total expenditure divided by accepted outputs

Human correction

Work required before publication

·····

Consumer watermarks may require additional disclosure when the scene could be mistaken for evidence.

A visible Grok watermark communicates that the asset was generated through Grok, although it may not provide enough context when the scene depicts a public figure, political event, testimonial, product demonstration, medical situation, or apparent news event.

Fictional entertainment content may require little additional explanation, whereas realistic media presented in an evidentiary or commercial setting may need a clearer statement that the scene, action, speech, or endorsement was synthetically created.

The publication context, audience, platform rules, and applicable law determine whether the watermark alone provides sufficient disclosure.

The watermark should remain visible rather than being cropped, covered, or altered.

·····

Commercial publication requires review beyond technical generation success.

A company should verify that it owns or has permission to process every reference image, brand asset, photograph, product design, and likeness included in the workflow.

The completed media should be inspected for accidental third-party logos, recognizable protected characters, inaccurate package text, unsupported product claims, misleading endorsements, and real-person resemblance.

Generated dialogue and captions require manual verification because one incorrect price, ingredient, warning, contractual statement, or legal claim can make an otherwise polished asset unusable.

The approval record should identify who accepted the prompt, references, generated output, edits, legal treatment, and final publication file.

·····

Current limitations determine where Grok Imagine fits within a production pipeline.

Consumer and API capabilities remain different, preventing API-level 1080p control from being promised to every paid subscriber.

The shared weekly allowance does not translate into one fixed number of images or videos, while high-resolution video consumes substantially more capacity than ordinary text interaction.

Consumer 720p requests may fall back to 480p, video duration remains capped at fifteen seconds, editing accepts source clips of approximately 8.7 seconds, and reference-to-video remains limited to 720p.

Image editing accepts only three source references, reference audio remains restricted, output URLs expire, and safety moderation continues after NSFW settings are enabled.

........

Current Grok Imagine limitations and their production consequences.

Limitation

Production consequence

Consumer and API features differ

Subscription users do not automatically receive developer controls

No fixed public generation count

Paid capacity varies with shared weekly usage

High-quality video consumes more allowance

Frequent creation can exhaust the weekly pool

Consumer 720p may fall back to 480p

Resolution can become inconsistent

Maximum video duration is 15 seconds

Longer sequences require several clips

Video editing accepts clips near 8.7 seconds

Longer source videos require another workflow

Reference-to-video is capped at 720p

The full 1080p mode is unavailable

Video editing is capped at 720p

Higher-resolution sources may be reduced

Image editing accepts three references

Complex compositions require staged generation

Reference audio has restricted eligibility

Ordinary users cannot assume access

Temporary output URLs expire

Media must be downloaded promptly

Consumer watermark cannot be removed

Paid access does not create a clean consumer output

NSFW does not disable moderation

Prohibited categories remain blocked

Exact moderation logic is unpublished

Borderline requests may behave unpredictably

Successful generation does not establish rights

Consent and intellectual-property review remain necessary

Repeated attempts multiply costs

Approved-asset cost exceeds the advertised output rate

·····

Grok Imagine operates most effectively when consumer experimentation and production controls remain separate.

The consumer product provides a conversational environment for generating, editing, restyling, and animating media without requiring structured development work.

Free access suits limited exploration, while SuperGrok and related paid plans increase the shared weekly capacity available across Grok products without guaranteeing a permanent number of visual generations.

The APIs provide the controlled route for applications requiring explicit models, 1K or 2K image output, defined aspect ratios, reference inputs, video durations, 480p through 1080p generation, moderation results, customer-managed storage, and predictable per-output pricing.

Safety restrictions remain independent from payment, allowing xAI to offer mature creative features while continuing to block sexualized minors, non-consensual intimate media, real-person nudification, deceptive impersonation, forged evidence, watermark removal, and attempts to bypass moderation.

A production team can therefore use the standard image model for concept development, move approved compositions into 2K Quality Mode, animate a verified still through Video 1.5, test lower resolutions before final rendering, inspect moderation status, download temporary outputs, and complete legal, visual, audio, brand, and safety review before publication.

·····

FOLLOW US FOR MORE.

·····

DATA STUDIOS

·····

·····DarLink AI Roleplay and Memory: character continuity, personalization, scenarios, and paid features

bottom of page