Grok can now stitch seven photos into a 30 second AI video
- Sugar Honey

- 7 days ago
- 2 min read
xAI just made AI video a genuine content tool instead of a novelty, and stock footage sites should be paying attention.
Grok Imagine's latest update lets Premium users feed in up to seven reference images, characters, locations, objects, whatever you've got, and generate a single cohesive 30 second video that keeps everything consistent across the whole clip. Faces, clothing, proportions and environments all stay locked in place frame to frame rather than warping the way early AI video famously did.
The engine behind it, xAI's Aurora architecture, predicts image tokens sequentially rather than generating the whole clip in one pass, which is reportedly what gives it tighter control over consistency. It's live now across iOS, Android and web for anyone with Grok Imagine access, alongside a 70 per cent price drop that makes experimenting with it a lot less painful.
Here's why seven reference images actually matters more than the headline video length. Most AI video tools still struggle to keep a single character looking like themselves for more than a few seconds. Being able to lock in a character from one photo, a location from another and props from several more, then have the model composite all of it into one scene, is the difference between a fun toy and something a creator could actually build a campaign around.
That's the real shift worth watching. AI video has spent two years being impressive in short bursts and unreliable at anything requiring consistency. Multi image reference tools like this are the bridge between "look what AI can do" demos and genuinely usable production tools for people who need the same character or product to show up correctly across an entire video.
If you make content for a living, this is worth an actual hands on test before you write it off as another gimmick. The gap between AI video hype and AI video utility just got noticeably smaller.




Comments