Contents

146 / 153

The Research Frontier

The state of the art, mid-2026

Chapter 145

11 min read

Reviewed v78 · August 2026

Everything above is the shape of the research frontier. This section is the snapshot of where that frontier actually sits as of this writing, the models released in the last thirty days, the ones leaked and not yet officially announced, the ones rumored for the next few weeks, and the trend-level observations that tie them together. This section will go stale faster than anything else in the document, but the goal here is not to be evergreen. The goal is to tell you exactly what is happening at the edge of the field right now, so you can calibrate everything else you read against real current state.

01

The Sora shutdown and what it means

The biggest single event in generative imagery over the past thirty days is OpenAI's announcement, on March 24, 2026, that Sora is being discontinued. The full timeline of the shutdown (web access ending April 26, API ending September 24, the details of what OpenAI said and what they probably meant) is covered in the OpenAI Sora chapter, so we will not retell it here. What matters for the research frontier is the strategic signal the shutdown sends.

02

The current video frontier

Earlier in 2026, the working hierarchy in frontier video generation looked roughly like this. Google Veo 3.1 was the overall quality leader, with native 48 kHz audio, native 4K output, and the strongest cinematic camera work of any model. Kling 3.0, released February 4, 2026, introduced Multi-Shot Storyboard (defining a sequence of shots with individual prompts, camera angles, and continuity between them) and was the first model to ship native 4K as standard. Seedance 2.0, from ByteDance, is the character consistency leader and delivered strong unified joint audio-video generation, the capability Veo 3 introduced natively in May 2025. Runway Gen-4.5, Pika 2, and Luma Ray 2 sit behind the leaders but have the best creator-workflow integrations. Wan 2.6 is the strongest open-weight model. Hailuo 2.3, PixVerse V5.5, and LTXV 2 round out the production-ready tier. Out of these, four of the six major models now generate synchronized audio natively, up from zero as recently as early 2025.

The practical shift for operators is that most production teams no longer commit to a single model. They route between two or three depending on the scene type: Veo 3.1 for hero shots that need the best visual quality, Kling 3.0 for multi-shot sequences and human performance, Seedance 2.0 for character-driven scenes that need consistency across cuts, Wan 2.6 when cost matters more than quality ceiling. Adobe Firefly quietly integrated Kling in late March 2026, which is the clearest signal yet that the incumbent creative tooling companies are embracing rather than fighting the open video model ecosystem.

03

FLUX.2 and the image model frontier

On the image side, the current frontier is held by Black Forest Labs's FLUX.2 family, released in late November 2025 and rapidly iterated since. The family has four main variants: FLUX.2 [klein] is the fastest, released January 15, 2026 and optimized for interactive generation; FLUX.2 [dev] is the open-weight variant with a commercial license; FLUX.2 [flex] exposes a steps parameter that trades off quality for latency at inference time; FLUX.2 [pro] and FLUX.2 [max] are the hosted production tiers with 4-megapixel photorealistic output. FLUX.2 introduces multi-reference conditioning, up to ten reference images per generation, and the family was released alongside an open-sourced FLUX.2 VAE under Apache 2.0, which is unusual, the VAE is usually where labs hide their competitive advantage. A March 3, 2026 update doubled generation speed across the family. This is the model currently being referenced throughout the document as 'FLUX,' and the technical details of the FLUX.1 era (12 billion parameters, rectified flow, 2024 training) should be mentally upgraded to the FLUX.2 specifics (4MP output, multi-reference conditioning, JSON prompting support, the open-sourced VAE) wherever the document's narrative feels behind.

The other contenders at the image frontier are: Google's Imagen 4 Ultra, which Google released in three genuinely differentiated tiers (the first major lab to ship tiered image generation with real quality differences rather than just speed variants), targeted specifically at real-world photography with attention to subsurface scattering and specular highlights; Midjourney 7, still the aesthetic leader for stylized work; Ideogram V3 and Recraft V4 Pro, both leaders in design-specific work with legible text; and Qwen Image 2 Pro from Alibaba, which is the strongest open-weight challenger. The gap between 'AI-generated' and 'photographically real' that has been the headline of the field for three years has, by 2026, officially closed for straightforward photorealistic use cases. The remaining gaps are in edge cases, complex compositions, and the failure modes covered in the common-questions chapter.

04

The GPT-Image-2 launch

On April 4, 2026, three anonymous image models appeared on LM Arena under the codenames maskingtape-alpha, gaffertape-alpha, and packingtape-alpha. They were pulled within hours once the community identified them, but not before multiple developers tested and posted results. On April 21, 2026, OpenAI officially launched GPT-Image-2, confirming the leak models were early variants. GPT-Image-2's significance is architectural, not positional. It is autoregressive rather than diffusion-based, reasons about the image before generating, achieves roughly 99 percent text rendering accuracy, and outputs at native 2K resolution. The reasoning step is why it opened the widest lead the arena had ever recorded, but that lead is evidence of the shift, not the reason to care about it. DALL-E 2 and DALL-E 3 are being retired on May 12, 2026.

What made the models unusual, according to community testing, was the combination of world knowledge and text rendering. One tester found that packingtape-alpha correctly rendered the time displayed on a watch face, which Nano Banana Pro had failed at. Another found that in a first-person Minecraft scene set in Manhattan, maskingtape-alpha outperformed Nano Banana Pro and its own sibling variants. The yellow-filter color cast that had been a signature flaw of GPT-Image-1 appears to have been fixed. Photorealistic portraits produced by the models were described by multiple testers as indistinguishable from real photographs. The models still failed the Rubik's Cube reflection test, which has been a standard spatial reasoning stress test, so the frontier on spatial reasoning has not moved even though the frontier on realism has.

The architectural significance of GPT-Image-2 is that it is autoregressive, not diffusion-based, marking a fundamental shift away from the DALL-E lineage. OpenAI is now producing images the same way it produces text, token by token, with the full reasoning capability of the language model available during generation. This is the same pattern Google uses with Nano Banana, but GPT-Image-2's Arena dominance suggests OpenAI has executed it more effectively. The retirement of DALL-E 2 and DALL-E 3 on May 12, 2026 makes this transition official: OpenAI's image generation future is autoregressive, not diffusion.

05

The open-source surge

The most consequential trend of the past month is that open-weight models are catching up to, and in some specific cases beating, the proprietary frontier. The headline release was Qwen 3.5 Small from Alibaba, released March 1, 2026, a family of four natively multimodal models at 0.8, 2, 4, and 9 billion parameters, all under Apache 2.0. The 9B variant beats GPT-OSS 120B on several graduate-level reasoning benchmarks despite being an order of magnitude smaller. On video understanding specifically, the Qwen 3.5 Small 9B outperforms Gemini 2.5 Flash-Lite. This matters because video understanding is a prerequisite for video generation, and Alibaba now has a genuinely world-class open video-understanding model that anyone can download and fine-tune.

The March open-source wave also included: Helios, a 14-billion-parameter autoregressive diffusion video model from Peking University, ByteDance, and Canva, released under Apache 2.0, capable of generating up to 1,440 frames (approximately 60 seconds at 24 fps) at 19.5 frames per second on a single NVIDIA H100. HY-WorldPlay from Tencent Hunyuan, released March 8, 2026, which published the reinforcement-learning post-training code for building real-time interactive world models on top of HunyuanVideo at 24 fps. Spectrum, a CVPR 2026 paper introducing a training-free spectral diffusion feature forecaster using Chebyshev polynomials that achieves up to 4.79x speedup on FLUX.1 and 4.67x on Wan2.1-14B without quality loss. Google Gemma 4, released in mid-2026, an encoder-free multimodal model across vision, video, and audio, with the 12B variant runnable locally in roughly 16 gigabytes of RAM (less for the smaller and quantized builds). Netflix's first public model on Hugging Face, called VOID (Video Object and Interaction Deletion), released April 2, 2026.

HappyHorse-1.0 is the most interesting development of the past week. A mysterious video model appeared on Artificial Analysis's benchmark leaderboard around April 7 without identifying its affiliations, climbed to the top of blind-test rankings for both text-to-video and image-to-video generation, and was revealed on April 10 to be a project from Alibaba's ATH AI Innovation Unit, still under active development. Alibaba confirmed the disclosure to CNBC. HappyHorse is not yet released but its leaderboard performance suggests that Alibaba has something at or ahead of Veo 3.1 level quality that it plans to ship openly. If that holds up under broader testing, it would be the first time an open-weight video model genuinely matched the proprietary frontier, which would be a significant inflection point for the field.

06

What is rumored and coming soon

Meta has publicly committed to shipping two new models in the first half of 2026 under the Superintelligence Lab led by Alexandr Wang. The image-and-video model is codenamed Mango, the text model is codenamed Avocado. These are the first major model releases from Meta's reorganized AI unit, and they carry significant weight because Meta does not currently have a competitive generative imagery product despite owning the infrastructure and data. Mango's positioning has been described as aiming beyond simple image and video recognition toward models that reason, plan, and act, which is consistent with the world-model framing OpenAI has been pushing for Sora's successor. Realistic release window is April through June 2026.

OpenAI's next-generation base model is internally codenamed Spud. Pretraining completed on March 24, 2026, the same day the Sora shutdown was announced, which is almost certainly not a coincidence. Sam Altman has indicated a release window of 'a few weeks,' which points to April or May 2026. Spud is expected to ship as GPT-5.5 or GPT-6 depending on how its benchmarks land. The relevance to generative imagery is that OpenAI's current image and video models are tightly coupled to its base models, and a new base model historically means a new generation of image and video capabilities shortly after.

DeepSeek V4 is notable less for its capabilities than for its hardware story. DeepSeek, working with Huawei and Cambricon, is adapting V4 so that inference runs on Huawei Ascend 950PR chips, with training still relying on NVIDIA GPUs, part of a push to reduce dependence on the NVIDIA CUDA ecosystem in favor of Huawei's CANN architecture. Hundreds of thousands of chips have reportedly been ordered as of mid-2026. This follows a failed earlier attempt with the Ascend 910C. The implication is that the Chinese open-model labs are actively building a parallel hardware stack that does not depend on NVIDIA, and if DeepSeek V4 ships successfully on this stack, it will be the first frontier-class model trained entirely outside the American chip ecosystem. For the generative imagery field specifically, this does not yet matter, but it probably will within two years, because every structural shift in the language model world eventually shows up in the image and video model world.

07

The trend-level observations

Pulling back from the individual releases, five things have shifted at the frontier in the past thirty days. First, native audio generation has gone from experimental to standard, four of six major video models now generate synchronized audio, up from zero as recently as early 2025. Second, native 4K output has gone from aspirational to table-stakes for the video frontier, with Kling 3.0 and Veo 3.1 both shipping it and Sora's lack of it cited as a reason the model fell behind. Third, the major production teams have stopped committing to a single video model and started routing between two or three based on scene type, which changes the competitive dynamics for every vendor. Fourth, the open-weight ecosystem is catching up faster than anyone expected, with Wan, Qwen, Helios, HY-WorldPlay, and HappyHorse all shipping at levels that would have been considered proprietary-only six months ago. Fifth, the Arena-first anonymous launch pattern pioneered by Google with Nano Banana has become the standard way major labs de-risk a new model before official release, and OpenAI's GPT-Image-2 leak confirms this pattern is now industry-wide.

Check your understanding

pass: 5 of 7

Answer at least 5 of 7 correctly to unlock the next chapter.

  1. 1. What makes the Sora shutdown historically notable?

  2. 2. Why is GPT-Image-2 architecturally significant?

  3. 3. What remained unsolved even as GPT-Image-2 pushed photorealism forward?

  4. 4. Why does Qwen 3.5 Small's strong video understanding matter for video generation?

  5. 5. Why would HappyHorse-1.0 be a significant inflection point if its performance holds up?

  6. 6. What is unusual about the FLUX.2 release beyond its variants?

  7. 7. What is notable about DeepSeek V4, more than its capabilities?

The weekly briefing

Get the week's moves in your inbox.

A short, sourced digest of what actually moved across generative AI, every week. Free.

Free. One email a week, no spam, unsubscribe anytime. Prefer a reader? RSS.