From 3D capture to video generation
Luma Labs has one of the more unusual origin stories in this document. The company was founded in San Francisco around 2021 by Amit Jain (a former Apple engineer who worked on the camera and AR teams) and Alex Yu (a Berkeley researcher who worked on neural radiance fields). The original Luma product was not a video generator at all, it was a 3D capture tool. You could point your phone camera at an object or scene, walk around it for a minute, and Luma's app would reconstruct a 3D model of what you captured using a technique called Neural Radiance Fields (NeRF). NeRFs were one of the cool computer vision techniques of 2021-2022, and Luma was the most polished consumer implementation of them.
Then in 2024, Luma made an unexpected pivot into video generation with a product called Dream Machine. The architectural connection makes sense in retrospect, both NeRFs and video generation involve building neural networks that understand 3D structure and how it changes over time, but the pivot was nonetheless dramatic. Within a year, Luma had become one of the major brand names in AI video generation, raised about $70 million in funding, and built a credible competitor to the bigger labs.
The Ray model line
Luma's video models are called Ray (short for 'Ray tracing,' a reference to their NeRF heritage). The first version was Ray 1, integrated into Dream Machine, released in mid-2024. It was decent but not at the frontier.
Ray 2 was released in early 2025 and represented a major upgrade, Luma claims it was trained on 10 times the compute of Ray 1, and the difference shows. Ray 2 produces fast coherent motion, ultra-realistic detail, and physically plausible event sequences that feel substantively more 'alive' than its predecessor. The technical approach is described as a multi-modal architecture, meaning it handles text, image, and video inputs in a unified model, with strong physical understanding inherited from Luma's 3D heritage. There is also a Ray 2 Flash variant for faster, cheaper iteration.
Ray 3, released in 2025, introduced 'Hi-Fi Diffusion' (a higher-quality master mode that produces 4K HDR output), Draft Mode for rapid iteration, and significantly better character consistency across long generations. The current flagship, as of July 2026, is Ray3.2, announced on June 9, 2026. Ray3.2 leans hard into professional production: it adds frame-by-frame directability, so you can steer specific frames rather than only the prompt, and HDR generation with paired EXR outputs that drop cleanly into a color and compositing pipeline. A week later, on June 16, Luma added 'Skills,' repeatable agent workflows that let a proven sequence of steps be saved and rerun. Ray positions Luma as a serious competitor to Veo and Kling in the high-end professional market, with the EXR and camera-motion work aimed squarely at teams that finish in a real post pipeline.
One feature unique to Luma's models is their physically grounded camera motion. Because Ray inherits architectural ideas from NeRFs, it understands 3D scene geometry in a more principled way than models that are purely 2D-aware. When you ask for a dolly shot, the camera actually moves through what looks like a real 3D space, with correct parallax. When you ask for an orbit, the subject stays geometrically consistent. This is subtle but it makes Luma's videos feel more like footage shot with a physical camera and less like AI-generated motion.
Luma is available through Dream Machine's web interface, the iOS app, Amazon Bedrock (Luma was the first video model to be integrated into AWS's enterprise AI platform), and through partner platforms including Krea. Pricing is on the lower end of premium, comparable to Hailuo, less than Veo or Sora.
The world-model turn, and the Saudi money behind it
Luma no longer describes itself as a video-generation company. Its stated mission is to build unified general intelligence that can generate, understand, and operate in the physical world, and the product line has followed that thesis: Ray3 was marketed as the first reasoning video model, emitting text and visual tokens to plan and judge its own output, and in March 2026 Luma launched Luma Agents on Uni-1, the first model in a Unified Intelligence family trained jointly on audio, video, image, language, and spatial reasoning. The pitch is a shift from prompting to directing, with an agent that keeps context, self-evaluates, and can even orchestrate rival models like Veo and Seedream inside its own workflow.
That ambition is expensive, which is where the money comes in. In November 2025 Luma raised a 900-million-dollar Series C led by HUMAIN, a company of Saudi Arabia's Public Investment Fund, at roughly a four-billion-dollar valuation, and tied it to becoming an anchor customer of HUMAIN's Project Halo, a reported two-gigawatt AI supercluster in Saudi Arabia. It is one of the largest AI raises of 2025 and it hard-wires Luma's compute future, and some of its optics, to the Gulf. The company's research is real too: its Inductive Moment Matching paper argues for breaking past diffusion's few-step ceiling, though whether that method powers the shipping Ray models is not disclosed.
Strengths and weaknesses
Luma's strengths are cinematic camera motion, which filmmakers responded to from the first Dream Machine launch, and an unusually professional output path: native HDR and 16-bit EXR export aimed at real color-grading and compositing workflows, which almost no consumer video model offers. It gives fine creative control through multi-keyframe direction, extend, loop, and reframe, plus a Draft Mode for fast, cheap iteration before a final render, and it carries a deep 3D and NeRF heritage that informs its physics-aware, world-model positioning. The HUMAIN capital and compute give it room to train at scale.
The weaknesses start with that reliance on self-reported benchmarks, and continue into a fast-moving, confusingly versioned lineup: Dream Machine, Ray2, Ray3, Ray3.14, Ray3.2, Ray3 Modify, Photon, UNI-1.1, and Luma Agents, with Dream Machine and Ray2 already deprecated, which can strand workflows built on a specific model. Pricing is fragmented across app tiers, agent tiers, and a separate API wallet whose credits do not transfer, it competes with and even embeds far better-funded rivals, and the Saudi PIF anchoring carries real dependency and optics risk.
Getting the best out of Luma
You reach the models through the Luma app and the Dream Machine web surface, the developer API with credit billing, and the newer Luma Agents front end for orchestrated, multi-step work. Watch the credit structure: app credits and API credits are reported not to transfer, annual billing saves meaningfully, and top-up packs bill separately. It is at its best on cinematic camera motion, HDR and EXR outputs for professional pipelines, keyframe-directed shots, and fast iteration in Draft Mode, and, like the whole category, at its weakest on long-duration consistency, precise on-screen text and hands, and any promise that the vendor's speed and quality numbers hold for your specific job.
The way to work it is to direct rather than prompt: describe the shot like a cinematographer with subject, action, camera move, lens, lighting, and mood, then use keyframes, extend, and loop for control instead of one giant prompt, and start from an image for identity and style consistency. Iterate in Draft Mode, keep individual clips short, and stitch for length. If you need a finished, gradable deliverable, Luma's HDR and EXR path is the actual reason to choose it over a flashier rival.