Specialized Models: Editing, Lipsync, 3D, and Upscaling
3D generation models
Chapter 62
2 min read
Reviewed v78 · August 2026
3D generation models
Generating 3D content is one of the harder frontiers in generative AI. A 3D model needs to be coherent from every angle, not just from one viewpoint, and the underlying representations are more complex than 2D images. There has been steady progress in this area, and several models are now usable for production work.
The most successful 3D generation model line at the moment. Tencent has released multiple versions of Hunyuan 3D as open-source models. The current generation generates textured 3D meshes from either text prompts or single reference images. The models are accessible through ComfyUIComfyUIThe dominant node-based interface for running diffusion models locally. Started by 'comfyanonymous' in January 2023, now stewarded by Comfy Org. (with native support), through Tencent's cloud APIAPI (Application Programming Interface)The programmatic endpoint you call to run a hosted model from your own code instead of a web interface., and through aggregatorAggregatorA tool or platform that wraps many image and video models behind one interface, so you switch models without switching apps. platforms. The quality is good enough for game asset prototyping, 3D-printed sculpture concepts, and product visualization mockups, though not yet at the level where you can use the outputs directly in a AAA video game without significant cleanup.
TRELLIS (Microsoft Research)
An open-source 3D generation model from Microsoft Research, released in late 2024. TRELLIS uses a 'structured 3D latent' representation and can generate meshes from text or image prompts. It is particularly good at preserving fine details from the input image when doing image-to-3D work.
InstantMesh, Wonder3D, and the academic models
There is a steady stream of academic 3D generation models, InstantMesh, Wonder3D, MVDream, and many others, that demonstrate techniques but are not yet polished enough to be production tools. Most are accessible through Hugging FaceHugging FaceThe dominant repository for open-source AI models, including all the open-weight image and video models discussed in this document. Spaces or as ComfyUI custom nodes. They are useful as research baselines and as starting points for custom workflows but are not yet what most working creators reach for.
3D generation is the area of generative AI most likely to see a major breakthrough in the next year. The current quality is good but not great, the architectures are still evolving, and several well-funded labs are working specifically on the problem. By the next major revision of this document, the 3D section may look quite different.
The frontier here has moved fast enough to change what the tools are for. Tencent's Hunyuan 3D now outputs meshes with physically based materials, automatic retopologyRetopologyRebuilding a generated 3D mesh with clean, efficient geometry so it can be animated and used in a real production pipeline., and part-splitting, which is the difference between a blob you have to rebuild and an asset a game team can actually drop into an engine. Microsoft's TRELLIS added LoRALoRA (Low-Rank Adaptation)A fine-tuning technique that adapts a base model to a specific style or concept using a small additional file. The standard way to customize open-source models. adapters, so a studio can bias generation toward its own asset style, and the fast image-to-meshMeshThe 3D geometry of an object, a network of vertices and faces, the core output of a 3D generation model. generators Meshy, Tripo, and Rodin cover the quick-draft end where speed beats fidelity. The reason to watch this category is structural: 3D is the one part of generative media where the output plugs directly into an existing, valuable pipeline, games, AR, virtual productionVirtual productionShooting live action against LED walls or in-camera digital backgrounds instead of a green screen, often using AI-generated environments., and product visualization, which means demand is real and specific rather than speculative.
Check your understanding
pass: 5 of 7
Answer at least 5 of 7 correctly to unlock the next chapter.
1. Why is generating 3D content harder than generating 2D images?
2. Which is described as the most successful 3D generation model line at the moment?
3. What inputs can Hunyuan 3D generate a textured mesh from?
4. What is the stated limitation of current 3D generation quality?
5. What representation does Microsoft's TRELLIS use?
6. What is TRELLIS particularly good at?
7. How are academic models like InstantMesh and Wonder3D characterized?
The weekly briefing
Get the week's moves in your inbox.
A short, sourced digest of what actually moved across generative AI, every week. Free.
Free. One email a week, no spam, unsubscribe anytime. Prefer a reader? RSS.