Gaussian Splatting Takes Flight: From Real-time Avatars to 4D Worlds and Beyond
Latest 32 papers on gaussian splatting: Aug. 1, 2026
Prepare for liftoff! Gaussian Splatting (3DGS) continues its meteoric rise, evolving from a fast rendering technique into a versatile backbone for some of the most exciting advancements in AI and 3D. This new wave of research pushes the boundaries of photorealism, efficiency, and real-time interaction, paving the way for truly immersive digital experiences. Let’s dive into the recent breakthroughs that are shaping the future of 3D content generation.
The Big Idea(s) & Core Innovations
At its heart, 3DGS excels at representing complex scenes with remarkable fidelity and speed. The core innovations explored in these papers revolve around extending this power to dynamic content, making it more compact, enabling intelligent interaction, and pushing its applicability into diverse fields.
One major theme is the real-time generation and animation of human avatars. The KAIST UVR Lab and KAIST KI-ITC ARRC, in their paper “S-Avatar: Diffusion-Guided Gaussian Head Avatars from a Single Image”, introduce a framework that generates animatable 3D head avatars from a single image using diffusion models and the FLAME parametric model. Their key insight is that combining diffusion priors with FLAME constraints dramatically improves 3D consistency, especially for unseen viewpoints, making single-shot avatar reconstruction more robust. Similarly, the University of Technology Sydney, Australia, presents “Split and Drive: Dual-Axis Disentanglement for Real-Time Gaussian Head Avatars”. This work champions a novel dual-axis disentanglement strategy, moving the driving computation inside the model and using specialized Gaussian branches for static, dynamic, and oral regions. This eliminates external tracker dependencies and drastically improves fidelity in complex areas like the mouth interior, as shown by their superior visual quality and expression accuracy. Complementing this, “FlexiAvatar: Unified 3D Gaussian Human Avatars Under Arbitrary Body Visibility” from UNIST, South Korea, addresses the challenge of partial visibility. Their visibility-aware optimization ensures that only observed body regions are updated, preventing ghost artifacts and reducing memory footprint, a crucial insight for robust avatar reconstruction from varied inputs.
Dynamic hair is also getting a significant upgrade. ETH Zürich, Switzerland, and Max Planck Institute for Intelligent Systems, Tübingen, Germany, introduce “Head Avatars with Dynamic Explicit Hair”, which leverages structured 3D Gaussian Splatting with an LSTM-based temporal motion model. This model conditions hair deformations on head angular velocity and gravity, resulting in physically plausible and temporally coherent hair dynamics, a significant step beyond previous unstructured methods.
The push for efficient and compact representations is equally strong. Xiaomi Technology Netherlands B.V. proposes “TSOG: A Format For Temporally And Spatially Ordered Gaussians”, a new file format for 4DGS content. By converting temporal evolution into index-aligned image data, TSOG achieves over 90% file size reduction for dynamic scenes, making real-time 4D streaming a reality. “3DGBGS: 3D Granular Ball Gaussian Splatting for Compact Novel View Synthesis” from Chongqing University of Posts and Telecommunications, China, further optimizes compactness by organizing SfM point clouds into adaptive granular balls. This method uses Granular Ball Anchor Initialization and Scale Prior to reduce anchor redundancy by over 37% while maintaining quality. In the realm of generalizable compression, Shanghai Jiao Tong University introduces “GenSplatCodec: Feed-Forward Gaussian Splatting Compression via One-Step Diffusion”. This codec reformulates low-bitrate compression as geometry-guided generative decoding, achieving a remarkable 17x storage reduction by separating structural coding from detail synthesis. Similarly, “ATSplat: Compact Feed-forward 3D Gaussian Splatting with Adaptive Token Expansion” from Yonsei University, South Korea, shows that high-quality feed-forward 3DGS comes from smart capacity allocation rather than dense sampling. Their adaptive token expansion achieves state-of-the-art rendering with 5.7x fewer Gaussians, rendering at an astonishing 1136 FPS.
Beyond visual realism, semantic understanding and real-world application are expanding. Peking University, and InkMind.AI unveil “ZeroSplat: Generalized Referring Segmentation in 3D Gaussian Splatting”, a training-free framework for generalized referring segmentation that can identify zero, one, or multiple targets in a 3D scene from text prompts. Their key insight is lifting 2D Vision-Language Model priors into 3D space via multi-view geometric constraints for intrinsic point-level understanding without per-scene optimization. For robotics and embodied AI, AMD AIG Team and Beijing Institute of Technology present “Real2Sim2Real for Vision-Language-Action Manipulation: An AMD ROCm-Based Pipeline”. This work demonstrates an AMD-accelerated pipeline for Sim-to-Real manipulation, combining 3DGS scene reconstruction with physics simulation to generate photorealistic training data. This proves VLA training and deployment can thrive outside the CUDA-locked ecosystem. Furthermore, Harbin Institute of Technology, China, presents “4DHumanDiff: Direct Text-to-4DGS Generation for Consistent 360-Degree Dynamic Humans”, a diffusion framework that directly generates 4DGS-based dynamic humans from text prompts. This bypasses computationally expensive video-first pipelines, yielding superior temporal and multi-view consistency at 10x faster inference.
Advanced editing and physics integration are also on the rise. “TOM-GS: Editable Video Representation via Temporal Opacity Modulation of Static 3D Gaussians” by University of Warsaw, Poland, introduces an editable video representation where static 3D Gaussians model scene dynamics through learnable temporal opacity. This allows for compatibility with standard 3D editing tools and physics engines. Fudan University, China, introduces “Inter-Reflective Gaussian Splatting for Robust and Efficient Inverse Rendering”, or IRGS++, which uses differentiable 2D Gaussian ray tracing for physically faithful inter-reflective rendering. This extends Gaussian inverse rendering to glossy, specular, and metallic materials, a significant leap in material fidelity. Donghua University, China, and Shanghai Artificial Intelligence Laboratory’s “Fashion-3DLR: A Controllable 3D Garment Generation Using Pairwise Fashion Elements for Intelligent Design” uses physics-integrated 3D Gaussians with a Material Point Method (MPM) solver to simulate fabric-specific dynamics (cotton, silk, wool, nylon) in generated 3D garments. In terms of scene editing, University of Stuttgart, Germany, contributes “3D-GIMP: When 3D Gaussian Inpainting Meets PatchMatch”, a hybrid approach for high-fidelity object removal that combines single-view generative inpainting with 3D-aware PatchMatch propagation, achieving comparable quality to diffusion methods at significantly faster speeds. Seoul National University, South Korea, introduces “Look Before You Edit: Attention-Guided Camera Placement and Multi-View Alignment for 3D Gaussian Splatting Editing”, tackling text-driven 3DGS editing by using attention-guided camera placement and multi-view attention alignment. This allows for highly localized, consistent edits with as few as 5 cameras, improving efficiency by up to 7x.
Even scientific visualization and wireless communication benefit from 3DGS. Temple University, USA, presents “3D Gaussian Splatting for Scientific Particle Data Compression and Rendering”, showing that 3DGS can compress and render large-scale scientific particle data (e.g., cosmological simulations) by 65-290x while outperforming traditional lossy compressors in visual fidelity. For wireless applications, “CORF-GS: Real-Time Wireless Radiance Field Reconstruction via Coupled Optical-RF Gaussian Splatting” and “Construction and Dynamic Update of Channel Gain Maps via 3D Gaussian Splatting” by various affiliations including Xiaomi and The Chinese University of Hong Kong (Shenzhen) show how 3DGS can reconstruct and dynamically update wireless radiance fields and channel gain maps. This enables real-time channel modeling for 6G networks by explicitly coupling optical and RF modalities or using physics-informed Gaussian primitives, leading to faster reconstruction and efficient updates.
Finally, for robust reconstruction and animation of dynamic scenes, Huazhong University of Science and Technology, China, and Nanyang Technological University, Singapore, present “Deblur-Avatar: Animatable Avatars from Motion-Blurred Monocular Videos”. This work tackles the challenge of motion blur in monocular videos by explicitly modeling human motion trajectories and jointly optimizing trajectories and 3D Gaussians for sharp avatar reconstruction. Nanjing University of Aeronautics and Astronautics, China, contributes “GrainGS: Gradient-Decoupled Gaussian Splatting for Efficient Dynamic Novel View Synthesis”, which uses a hierarchical anchor scaffold with a stop-gradient mechanism to prevent gradient interference between geometry and deformation, achieving state-of-the-art quality, real-time rendering (435.6 FPS), and compact storage. University of Illinois Urbana-Champaign, USA, and Waymo introduce “AniGS: Bridging Rendering and Diffusion Prior for 3D Scene Animation”, which animates static 3DGS scenes with subtle, distributed ambient dynamics (like vegetation motion) using an iterative dataset-model update strategy and video diffusion priors. This allows for free-viewpoint video rendering of animated large-scale scenes without ground-truth dynamic supervision.
For quality assessment, “SpatialQ: Spatial-Aware Multimodal Reasoning for 3D Gaussian Splatting Quality Assessment” by anonymous authors introduces a multimodal framework that combines 3D-aware representation learning with a grounded multimodal language model (MLLM). This enables interpretable degradation analysis, distinguishing between appearance and geometric consistency, crucial for robust quality metrics.
Under the Hood: Models, Datasets, & Benchmarks
The innovations in these papers are often underpinned by specialized models, novel datasets, and rigorous benchmarks. Here’s a look at some of the significant resources:
- S-Avatar: Leverages FLAME parametric model for facial geometry and diffusion models for initial splat generation. Code available at https://github.com/hailsong/savatar.
- TSOG: Extends the SOG framework with timeline attributes and temporal parameterization. Evaluated on SIGGRAPH Asia 2025 Volumetric Video Challenge Dataset and Neural 3D Dataset.
- Split and Drive: Uses an internal motion encoder and three specialized Gaussian branches (Dynamic, Static, Mouth Interior) for disentanglement.
- 4DHumanDiff: Constructs a large-scale text-to-4DGS dataset (60,000 pairs) and uses a specialized 3D U-Net with temporal attention.
- StructureGS: Integrates Oriented Bounding Boxes (OBBs) for structural guidance, evaluated on datasets like PARIS (https://github.com/Jianghongbb/PARIS).
- SpatialQ: Develops a spatial quality encoder and integrates a grounded MLLM (Qwen2.5-VL 7B). Achieves state-of-the-art on 3DGS-IEval-15K dataset (https://github.com/aim-uofa/3DGS-IEval-15K).
- 3DGBGS: Introduces Granular Ball Anchor Initialization (GBAI) and Granular Ball Scale Prior (GBSP), evaluated on Mip-NeRF 360, Tanks&Temples, Deep Blending, BungeeNeRF.
- CORF-GS: Uses a unified Gaussian representation with shared geometry and modality-specific appearance. Evaluated on a custom Optical-RF dataset from NIST lobby (via Blender and Sionna RT).
- SONG: Combines 3DGS rendering with LLM-generated semantic pedestrian behaviors and Kimodo-synthesized natural full-body motions. Benchmarked with SONG-Bench and integrates with Isaac Lab. Code available at https://github.com/AISimulation/SONG.
- GenSplatCodec: Employs a dual-stream architecture (structural and appearance) and one-step diffusion. Evaluated on DL3DV and RealEstate10K datasets.
- DynHair: Leverages structured hair strands with an LSTM encoder and FiLM conditioning. Data and code available at https://dynhair.is.tue.mpg.de/.
- Fashion-3DLR: Features a Garment Feature Fusion Diffusion Transformer (GFF-DiT) and integrates with 3D Gaussian-based physical simulation (Material Point Method). Utilizes SewFactory dataset.
- Real2Sim2Real: Uses SmolVLA model, Genesis physics engine (https://genesis-embodied-ai.github.io/), and an AMD ROCm-based pipeline. Public code available at https://github.com/AMD-AIM/Physical_AI_Challenge.
- ParticleGS: Introduces VizMapper (a lightweight neural network) for adaptive visualization parameters. Evaluated on HACC N-body cosmological simulation and FIRE-2 public data (https://github.com/BoJiang03/ParticleGS).
- Meshless Domain Randomization: Modulates Spherical Harmonics coefficients and applies 3D spatial noise to 3DGS parameters. Uses a UnityGaussianSplatting plugin (https://github.com/aras-p/UnityGaussianSplatting).
- Inter-Reflective Gaussian Splatting: Uses differentiable 2D Gaussian ray tracing and metallic-aware BRDF modeling. Evaluated on TensoIR, GlossySynthetic, RefReal, GlossyReal, Stanford-ORB datasets.
- TOM-GS: Modulates opacity with learnable temporal mean and scale parameters. Uses DAVIS dataset and AnyCam (https://github.com/anycam/anycam).
- Deblur-Avatar: Models human motion trajectories and uses pose-dependent fusion. Evaluated on ZJU-MoCap-Blur and Real-Human-Blur datasets (https://github.com/xianrui-luo/deblur-avatar).
- Visual Relocalization: Combines MVS depth/normal supervision with LiDAR-guided Chamfer loss. Evaluated on DLR S3LI Vulcano Dataset (https://github.com/DLR-RM/multimodal-gsplat-relocalization).
- GrainGS: Employs a hierarchical anchor scaffold with per-Gaussian deformation and stop-gradient isolation. Benchmarked on D-NeRF and DG-Mesh datasets.
- GLAM-SLAM: Uses a decoupled ORB-SLAM2 frontend with a parallelized 3D Gaussian mapping backend, including a flow-guided densification module and localized MLPs. Evaluated on KITTI, Oxford RobotCar, Málaga datasets (https://github.com/pmermigkas/GLAM-SLAM).
- Construction and Dynamic Update of Channel Gain Maps: Decomposes channel gain into distance-dependent attenuation, path transmittance, and effective scattering, represented by Gaussian primitives with physics-informed features. Uses OpenStreetMap and Sionna-RT.
- SubSplat: Features a Sub-pixel Gaussian Reparameterizer (SPGR) and deformable attention-based feature aggregation. Evaluated on RealEstate10K and ACID datasets.
- 3D-GIMP: Combines generative inpainting with 3D-aware PatchMatch propagation and Poisson-based depth completion. Uses IMFine, 360-USID, Mip-NeRF 360 datasets.
- RealVDeblur: A diffusion-based framework that uses a large-scale physically grounded blur synthesis pipeline (OmniBlur with ~2,000 3DGS scenes) and a pre-trained video diffusion prior.
- ATSplat: Built around adaptive 3D anchor tokens and an Adaptive Token Expansion module. Resources at https://join16.github.io/page-atsplat.
- MR-Compare: Uses a two-stage coarse-to-fine registration pipeline (TEASER++ for coarse, V-GICP for fine) with a 3D Slider mechanism. Evaluated with Replica dataset and Meta Quest 3. Code at https://github.com/changruizhu96/MR-Compare.
- Look Before You Edit: Uses Attention-Guided Editing Camera Placement (ACP) and Multi-View Attention Alignment (MAA).
- FlexiAvatar: Employs an occlusion-robust SMPL-X tracking pipeline, part-specific residual refinement, and diffusion-based generative texture completion. Evaluated on NeuMan, ZJU-MoCap, WildAvatar, TalkShow, INSTA datasets. Project page: https://yihalem1.github.io/FlexiAvatar/.
- ZeroSplat: Introduces GR-LERF and GR-ScanNet benchmarks for evaluation. Project page: https://inkmind-ai.github.io/ZeroSplat.
- AniGS: Uses LTX-Video diffusion model for motion supervision and SAM2 segmentation model. Evaluated on DL3DV benchmark dataset.
- ECoNGS: Hybrid neural-explicit Gaussian splatting for volume visualization. Uses multi-scene joint optimization and neural entropy coding. Code available at https://github.com/TouKaienn/ECoNGS.
Impact & The Road Ahead
The impact of these advancements is monumental. We are rapidly moving towards a future where high-fidelity 3D content, previously requiring specialized expertise and expensive hardware, can be generated, animated, and interacted with in real-time from simple inputs. This democratizes 3D creation, enabling new possibilities across various industries:
- AR/VR and Gaming: Photorealistic, animatable avatars and dynamic scene generation will transform virtual worlds, making them more immersive and interactive. Imagine creating a personalized avatar from a single selfie or exploring dynamic environments with physically accurate hair and clothing.
- Robotics and Embodied AI: The ability to quickly generate synthetic data that bridges the sim-to-real gap, combined with semantic 3D understanding, will accelerate the development of more intelligent and robust robots capable of navigating complex, dynamic real-world environments.
- Digital Fashion and E-commerce: Real-time 3D garment generation with physics-based simulation will revolutionize online shopping and virtual try-ons, offering unparalleled realism.
- Scientific Visualization: Compressing and rendering massive scientific datasets with visual fidelity will enable researchers to interactively explore complex simulations, unlocking new insights.
- Wireless Communication: Real-time wireless channel modeling promises more efficient and reliable 6G networks, adapting to dynamic environments and optimizing signal propagation.
The road ahead for Gaussian Splatting is incredibly exciting. Key areas for future research include further improving the compactness and generalizability of 4D representations, pushing the boundaries of real-time photorealistic editing, and integrating even more sophisticated physics and semantic understanding directly into the Gaussian primitive properties. As the field continues to evolve at this rapid pace, we can expect to see 3DGS become an indispensable tool for bridging the physical and digital worlds, creating experiences that were once confined to science fiction.
Share this content:
Discover more from SciPapermill
Subscribe to get the latest posts sent to your email.
Post Comment