People and story
Midjourney is the strangest major company in this entire document, and the strangest in instructive ways. It was founded in 2022 by David Holz, who had previously co-founded Leap Motion, a hand-tracking sensor company that pioneered gestural interfaces before the VR boom made that capability commercially relevant. Holz funded Midjourney himself out of his Leap Motion proceeds and has famously refused to take outside venture capital. The company has never disclosed its size, but estimates put it at around 40 employees as of 2024, growing to an estimated 60-80 by 2026, generating hundreds of millions of dollars in annual subscription revenue. Midjourney does not have a Chief Financial Officer, a marketing team in any conventional sense, a sales team, or most of the other organizational apparatus you would expect from a company of its revenue. It has a community of millions of paying users who interact with the product through Discord and a web interface, and a founder who acts as primary researcher, primary product designer, and primary voice of the company.
Midjourney's product philosophy is unique and worth understanding before evaluating the models. Where other image labs optimize for photorealism, prompt accuracy, text rendering, or instruction following, Midjourney optimizes explicitly for aesthetic quality. The team has stated publicly that their goal is to make images that are beautiful, not necessarily images that match the prompt literally. This means Midjourney will often take liberties with your prompt to produce something more visually striking than what you literally asked for. For some users this is the entire point; for others it is a bug that makes Midjourney unusable for production work where the brief has to be followed exactly. Which camp you fall into determines whether Midjourney is the right tool for your work or whether you should use FLUX or Imagen instead. Neither answer is wrong.
The fact that Midjourney runs on Discord, where you generate images by typing slash-imagine commands into a chat channel that shows everyone else's outputs in real time, was an accidental masterstroke of community building. The Discord interface meant that every user was constantly exposed to other users' work, which fueled inspiration, learning, and the sense that you were part of an active creative community. Many designers and artists got their first exposure to AI imagery through Midjourney's Discord, and the community shapes the aesthetic sensibility of the model in subtle ways through the rating data that gets fed back into training. David Holz has been explicit that this community rating loop, where users can thumbs-up or thumbs-down each other's outputs and those ratings get used in model training, is one of Midjourney's core competitive advantages.
Architecture and core ideas
Midjourney is unusually opaque about its architecture. The company has never published a technical report, has never open-sourced any weights, has never released a research paper, and David Holz has explicitly stated that he considers the details of how Midjourney works to be a competitive advantage worth protecting. What we can infer from external analysis and occasional interviews is that the early versions (V1 through V3) used U-Net based diffusion architectures similar to the Disco Diffusion and VQGAN+CLIP families that were popular in 2021-2022. V4 appears to have been the transition to a more conventional latent diffusion model, and V5 onward are almost certainly diffusion transformer variants, though the exact configuration is unknown.
What distinguishes Midjourney architecturally is not the backbone but the training pipeline. The Midjourney team has been explicit that their model's aesthetic sensibility comes from the ratings data their users produce, billions of thumbs-up/thumbs-down signals on generated images, used as a reinforcement learning reward signal on top of the base diffusion training. This is structurally similar to the RLHF approach that tunes language models like ChatGPT, except applied to visual aesthetics rather than helpfulness. The effect is that Midjourney's outputs reflect the accumulated aesthetic preferences of millions of paying creative users, which produces a characteristic 'Midjourney look' that is more painterly, more dramatic, more cinematically composed, and more willing to depart from literal prompt interpretation than models trained purely on scraped web data.
Midjourney version history
V1 through V4 (2022)
The early versions established Midjourney's dreamlike, painterly aesthetic that became instantly recognizable. V1 was a curiosity, V2 was the first version that produced consistently interesting results, V3 added major coherence improvements, and V4, released in November 2022, was the version that shifted Midjourney from 'fun toy' to 'professional creative tool' in the minds of designers. V4 coincided with the original Stable Diffusion launch and established the competitive dynamic that has shaped the field since.
V5 and V5.2 (March-June 2023)
V5 was the version that achieved photorealism for the first time. Before V5, Midjourney images were obviously paintings or illustrations; V5 could produce outputs that were mistakable for photographs, which was a major capability jump for the field. V5.2 added image-to-image features and improved compositional control. The V5 era is when Midjourney became the dominant creative tool in the AI imagery space, ahead of DALL-E 2 and Stable Diffusion 1.5.
V6 and V6.1 (December 2023 - August 2024)
V6 introduced significantly better prompt following, which had been a long-running weakness of Midjourney. For the first time, you could write complex compositional prompts and have the model actually respect them. V6 also improved text rendering, though it still lagged behind Ideogram on this axis. V6.1 in August 2024 was a quality refinement release, iterating on V6's foundation rather than introducing new capabilities.
V7 (April 2025)
V7, released in alpha in April 2025, was Midjourney's first new major model architecture in nearly 18 months, and David Holz described it as a 'totally different architecture.' V7 introduced a personalization profile system, to use it at all, new users had to rate roughly 200 images so the model could learn their visual preferences and tune its outputs accordingly. This was a significant product decision, it made the first-use experience more demanding but produced dramatically better per-user results after the initial rating session. V7 also introduced Draft Mode (generating lower-quality images at roughly 10 times the speed for rapid iteration), voice input, and Midjourney's first foray into video generation with clips of up to 21 seconds.
V8 (late 2025, ongoing iteration)
V8 shipped in late 2025 as the production successor to V7, and a V8.1 point release, with faster and higher-resolution output on the same architecture, became the default in mid-2026. V8's headline improvements were better prompt adherence (finally closing most of the gap to FLUX.2 and Imagen 4 on instruction following), improved character consistency across generations, and expanded video capabilities including longer clip durations and better motion coherence. V8 is the first Midjourney version where the gap to the frontier general-purpose models has meaningfully closed rather than widened, which was not obvious it would happen given how far Midjourney had drifted from the general-purpose optimization path the rest of the field was on. V8 also retained Midjourney's signature aesthetic advantage, the characteristic painterly, cinematic look is still the best in the field for stylized work, even as the photorealism and prompt-following gaps narrowed.
Strengths and weaknesses
Midjourney's strengths are aesthetic quality on stylized and painterly work, cinematic composition, and the kinds of dreamlike or dramatic outputs that other models cannot match. For concept art, book covers, editorial illustration, mood boards, cinematic stills, and any use case where the image needs to feel striking rather than accurate, Midjourney is still the best single model available. The personalization profile system introduced in V7 also means that heavy Midjourney users get per-account aesthetic consistency that is hard to replicate in other tools. V8's improvements to prompt following have made Midjourney viable for commercial design work in a way it was not a year earlier.
Midjourney's weaknesses are ecosystem isolation, lack of API access, and unsuitability for certain professional workflows. There is no API, which means you cannot integrate Midjourney into your own product or use it in a workflow that chains multiple models together. This single limitation rules Midjourney out of most serious operator use cases covered in the Operator's Playbook, because operators need to route between many models and Midjourney refuses to participate in that routing. There are also no weights, no LoRAs, no fine-tuning, and no offline use. If your needs include brand consistency across a catalog, character persistence across a narrative, or any kind of custom training, Midjourney simply cannot do those things. Midjourney is a consumer creative tool, not an infrastructure component.
Strategic position
Midjourney occupies a strategic position that no other major AI lab has been able to replicate. It has no investors to answer to because it is fully founder-owned, which means it does not have to chase the scaling curves the rest of the field is racing along. It has a dedicated user base that pays for creative quality rather than for features or speed. It has a brand that is synonymous with AI art in the popular imagination, partly because of V5's viral moment in 2023 and partly because the Discord community dynamic made Midjourney the public face of generative imagery for most of its early years. And it has an aesthetic advantage that comes from its ratings-based training loop, an advantage that compounds as the community grows because more ratings produce better aesthetic tuning which attracts more creative users who produce more ratings.
The strategic risks are real, though. Midjourney's isolation from the broader ecosystem means it cannot benefit from the rapid innovation happening elsewhere, every new technique, every new base model, every new workflow pattern has to be reimplemented internally from scratch. The lack of API access means Midjourney misses the entire operator market, which is the fastest-growing segment of the generative imagery economy. And the founder-owned model means that when David Holz decides Midjourney is done, or decides to sell, or changes his mind about any of the core product decisions, there is no board or investor group to moderate the outcome. For now, Midjourney's willingness to stay small, stay focused, and stay weird has been exactly the right strategy for the aesthetic-focused creative market they serve. Whether that continues to work as the field matures and the operator market grows is one of the open questions about 2027 and beyond.
The studio lawsuits
Midjourney is the first generative image company Hollywood took to court, and it happened because the model will cheerfully produce recognizable copyrighted characters. On June 11, 2025, Disney and NBCUniversal jointly sued it in California, in a roughly 110-page complaint that called the service a bottomless pit of plagiarism and cited generations of Darth Vader, Elsa, the Minions, and Shrek. Midjourney responded in August that copyright law does not confer absolute control over how works are used. Then on September 4, 2025, Warner Bros. Discovery filed a separate suit over Superman, Batman, Bugs Bunny, and Rick and Morty, making three of the big five studios plaintiffs at once.
Getting the best out of Midjourney
Since August 2024 you no longer need Discord: the web app at midjourney.com shares your account and adds a real editor with inpainting, pan, and zoom, though the Discord bot still works. Plans run on a flat subscription, roughly ten dollars for Basic up to about a hundred and twenty for Mega, with higher tiers adding fast GPU hours, unlimited relax-mode generation, and, from Pro up, a Stealth mode that keeps your generations private, since lower tiers are public by default. The craft lives in the parameters: use --ar for aspect ratio, --stylize (or --s) to dial how hard Midjourney applies its own aesthetic, --sref to borrow a style from a reference image, the character and omni reference to hold a subject consistent, and RAW mode when you want it to stop imposing its default beauty.
It is best at striking single images, concept art, illustration, moodboards, and atmospheric or painterly scenes, and via the Niji models, anime. It is weakest at long precise in-image text, exact diagrammatic or interface layouts, and neutral, documentary, non-stylized imagery, which is the moment to reach for RAW and a low stylize value or to use a different tool entirely.
Where Midjourney is headed
David Holz has described a long arc from images to video to 3D to real-time, interactive, open-world simulation, sometimes analogized to a holodeck or a sandbox for making games and films, and the June 2025 image-to-video model, its V1, was framed as a deliberate step on that path. The company has also built a hardware team, hiring people with Leap Motion and Apple Vision Pro backgrounds, which echoes Holz's own past as the co-founder of the hand-tracking company Leap Motion. What makes this credible rather than hype is that Midjourney is famously bootstrapped and profitable, took no outside venture capital, and runs with a tiny team on flat subscription revenue, so it can chase a decade-long vision without answering to growth-at-all-costs investors. These are stated ambitions and early efforts, not shipped world-simulation or hardware, but the independence is real.