Gaussian Splatting Takes Flight: From Dynamic Scenes to Robotic Perception and Beyond
Latest 54 papers on gaussian splatting: Oct. 10, 2026
3D Gaussian Splatting (3DGS) has rapidly emerged as a powerful paradigm for novel view synthesis, offering impressive photorealism and real-time rendering capabilities. However, its initial formulation often grappled with dynamic scenes, sparse input data, geometric accuracy, and integration into broader AI systems. Recent research, as highlighted by a wave of innovative papers, is pushing the boundaries, transforming 3DGS into a versatile tool for dynamic content creation, robust robot perception, efficient compression, and even artistic expression.
The Big Idea(s) & Core Innovations
The overarching theme in these advancements is the quest for greater versatility and robustness in 3DGS, moving beyond static, dense captures. A key challenge addressed is handling dynamic scenes and sparse input, which often leads to artifacts. OuroWorld, from the National Yang Ming Chiao Tung University and Alaya Lab, tackles this head-on by transforming static 3DGS scenes into endlessly looping 3D cinemagraphs. Their OuroWorld: Bringing Any 3D World Alive as Diverse, Endlessly Looping 3D Cinemagraphs introduces an Inconsistency-Robust Periodic 4DGS with Fourier-series deformation to guarantee perfect looping and a Grounded Drift Field to absorb cross-view inconsistencies from generated multi-view videos. This enables diverse motion types, a significant leap from prior fluid-like dynamics.
For sparse-view scenarios, several papers introduce clever ways to augment data or guide reconstruction. Sparse-View 4D Gaussian Splatting via Spatiotemporal Priors and Generative Assistance by authors from Tsinghua University and JD.com, achieved first place in the SIGGRAPH Asia 2026 Volumetric Video Challenge by using region-adaptive spatial priors, motion-consistent temporal priors, and generative assistance from a diffusion model to reconstruct dynamic scenes from just six wide-baseline cameras. Similarly, UGOD: Uncertainty-Guided Opacity and Dropout for Sparse-View 3D Gaussian Splatting by a multi-institutional team including Manchester Metropolitan University and University of Surrey, leverages view-dependent uncertainty to modulate Gaussian opacity and drive soft dropout, reducing overfitting and improving compactness in sparse-view settings.
Another critical area of innovation focuses on improving geometric fidelity and enabling downstream tasks. PCAsplat: Gaussian Splatting with Local PCA Regularization from Universidade de São Paulo and Stanford University, introduces a differentiable local PCA framework that regularizes Gaussian neighborhoods directly in scene space. This encourages Gaussians to align with underlying surfaces, drastically reducing “floaters” and making splats suitable for tasks like segmentation and Poisson reconstruction. Reinforcing the importance of geometric priors, Prior-Driven Enhancements in 3D Gaussian Splatting: Normals and Depths Regularization by KakaoMobility, integrates surface normals and dense depth priors to refine 3DGS optimization, especially in challenging environments like reflective surfaces or low-texture regions.
Furthermore, researchers are exploring efficient representations and system-level optimizations. Budgeted-GS: Real-Time Large-Scale Gaussian Splatting via Factoring LOD by Haipeng Wang from Neusoft, transforms trained 3DGS models into a factoring tree (a multi-resolution hierarchy) for real-time, city-scale rendering across devices, even introducing “Budget-Centered Training” to train models directly at optimal capacity. For distributed training, ByteSplat: Efficient Distributed 3D Gaussian Splatting Training via Intra- and Inter-GPU communication reduction by Shanghai Jiao Tong University and IEIT SYSTEMS, drastically speeds up distributed training by fusing rasterization kernels and eliminating zero-gradient communication. Meanwhile, TileSkipper: Region-Adaptive Tile Pruning for 3D Gaussian Splatting from Arizona State University and Johns Hopkins University, optimizes rendering by compiling frozen 3DGS checkpoints into heterogeneous per-Gaussian tile pruning policies, achieving significant speedups with negligible quality loss.
Finally, the field is seeing a surge in semantic understanding, editing, and specialized applications. ManifoldSplat: Language-Guided Semantic Shape Editing of 3D Gaussian Head Avatars by Universitat Politècnica de Catalunya and Institut de Robòtica i Informàtica Industrial, enables language-guided semantic shape editing of animatable 3D Gaussian head avatars by performing edits within a structured FLAME manifold, preserving identity and animation. OpenSplatGraph: From Dense Semantic Maps to Structured Scene Graphs for Open-Vocabulary Robot Perception from Queensland University of Technology and CSIRO Robotics, bridges dense semantic mapping with structured object-centric reasoning for robots, creating persistent 3D scene graphs from Gaussian-based semantic maps. And for medical applications, FactorSplat: Appearance-Controllable Gaussian Proxies for Medical Volume Rendering by United Imaging Intelligence, allows region-specific intensity-to-RGBA transfer functions at inference time, enabling appearance changes in medical volumes without retraining.
Under the Hood: Models, Datasets, & Benchmarks
These papers frequently introduce or heavily rely on specialized models and datasets, pushing the capabilities and evaluation standards of 3DGS:
- OuroWorld (https://ouroworld.userwei.com): Utilizes vision-language models for dynamics inference, video diffusion models for reference video generation, and an Inconsistency-Robust Periodic 4DGS for looping. Evaluated on 39 reconstructed and generated scenes.
- LVS (https://arxiv.org/pdf/2610.12127): Uses a multiscale residual network and reuses rendered RGB-D images. Benchmarked on GS-render (Mip-NeRF 360 scenes) and Blender Barcelona Pavilion datasets.
- 2DGS-Planner (https://2dgs-planner.github.io/): Leverages rasterization as a geometric query interface for 2DGS maps. Uses public Real2Sim USDZ scenes and demonstrates real-world deployment on a humanoid robot.
- PAM-ToD (https://arxiv.org/pdf/2610.11572): A plug-in module with hierarchical multiplicative and additive terms. Introduces CARLA-ToD, a new benchmark with matched geometry and poses across different times of day.
- OX-NeRF (https://arxiv.org/pdf/2610.11547): Combines a cross-scene convolutional encoder with per-scene multi-resolution hash grids. Introduces Shells and Voids synthetic datasets for ultra-sparse X-ray reconstruction and evaluates on LIDC-IDRI.
- FlyMark (https://arxiv.org/pdf/2610.11364): Utilizes a published fruit fly connectome’s photoreceptor layout as a public, non-tunable carrier geometry for watermarking.
- Rendering-Free Lookahead (RFL) (https://arxiv.org/pdf/2610.11039): Employs a privileged teacher with 3DGS to render future views. Evaluated on E3VS-Bench and SceneSplat++ (99 indoor 3DGS scenes).
- PCAsplat (https://arxiv.org/pdf/2610.11011): Uses differentiable local PCA regularization. Evaluated on DTU, Tanks and Temples, and ScanNet++ datasets. Code will be released.
- GaussianBench (https://github.com/Dalu1234/GaussianBench): A physics-fidelity suite for 3DGS, evaluating seven broad profiles, including six external systems (PhysGaussian, GaussianFluent, OmniPhysGS, PhysDreamer, Physics3D, GASP). Introduces GaussianFlesh as a thermomechanical reference entrant.
- Sparse-GS2Mesh (https://arxiv.org/pdf/2610.04203): Combines 3DGS with stereo matching and 2DGS co-regularization. Evaluated on DTU and BlendedMVS datasets.
- GDSNet (https://arxiv.org/pdf/2610.10396): Represents crowds as continuous 2D Gaussian primitives with control-point-guided parameterization. Benchmarked on ShanghaiTech A and B, JHU-Crowd++, UCF-QNRF, and NWPU-Crowd datasets.
- DeltaSplat (https://arxiv.org/pdf/2610.09853): A lightweight iterative Gaussian refinement module using Plücker rays and rendered depth. Achieves SOTA on DL3DV dataset and generalizes to ScanNet++ and RE10K.
- SPLATIFY (https://arxiv.org/pdf/2610.09116): A multi-agent framework to convert 3DGS papers to code. Introduces SPLATIFY-BENCH, a 30-paper benchmark. Code, data, and implementations will be publicly released.
- DensiTok (https://hydragon.co.kr/DensiTok): A plug-in module using a Level-Adaptive Modulation VAE and a camera-conditioned Diffusion Transformer. Improves performance on RealEstate10K and DL3DV for backbones like AnySplat, WorldMirror, Depth Anything 3.
- GSCV (https://github.com/Qi-Yangsjtu/GSCV): Compresses 3DGS sequences using standard video codecs. Evaluated on MPEG GS compression dataset and compared against GSCodec Studio and GPCC v1.
- OpenSplatGraph (https://csiro-robotics.github.io/OpenSplatGraph): Constructs 3D scene graphs from open-vocabulary Gaussian semantic maps. Validated on standard 3D scene understanding benchmarks and real-world robotic experiments.
- SURGE (https://arxiv.org/pdf/2610.07472): A camera-sonar framework for underwater localization and reconstruction. Uses BlueROV2 and M750d forward-looking imaging sonar.
- MoonGS (https://github.com/InRobots/MoonBlender): Uses vision foundation models (DUSt3R, VGGT) and introduces MoonBlender, a synthetic lunar dataset.
- SteadySplats (https://arxiv.org/pdf/2610.05576): Introduces color variance regularization and history-based spatial resampling. Evaluated on Mip-NeRF 360 with a Vulkan-based renderer.
- VolS-GS (https://justin4ai.github.io/VolS-GS): Relightable GS with volumetric subsurface scattering. Evaluated on SSS-GS, NRHints, GS3 benchmarks.
- ManifoldSplat (https://a-canela.github.io/manifoldsplat/): Language-guided semantic shape editing for 3D Gaussian head avatars. Uses the FLAME manifold and a synthetic dataset of region-specific edits.
- Budgeted-GS (https://arxiv.org/pdf/2610.03162): Introduces Factoring LOD and Budget-Centered Training. Evaluated on Mip-NeRF 360 and MatrixCity datasets.
- ByteSplat (https://github.com/): For distributed 3DGS training. Evaluated on Mill-19, UrbanScene3D, MatrixCity, Mip-NeRF 360, DeepBlending, Tanks&Temples datasets.
- FactorSplat (https://gaozhongpai.github.io/FactorSplat/): A per-scene N-dimensional GS proxy for medical volume rendering. Evaluated on seven CT and MR scans.
- SCION (https://light.princeton.edu/SCION): Hierarchical compositional scene representation with reusable primitives. Demonstrates compact storage (1.2 MB).
- EvenSplat (https://arxiv.org/pdf/2610.01876): Decouples appearance from illumination. Introduces a nine-scene real-world captured dataset for evaluation.
- MEGA (https://arxiv.org/pdf/2610.01707): Extracts object-level, watertight meshes via Spatial Visual Distillation. Enables physical interactions.
- Affine-Aligned Atlas (https://arxiv.org/pdf/2610.01114): For video representation, uses frame-wise affine transforms. Integrates with D2GV, GSVR, STGV and evaluated on Big Buck Bunny, UVG.
- TRACE (https://arxiv.org/pdf/2610.00822): Privacy-preserving next-best-view selection for multi-robot systems. Uses Habitat-Sim and Gibson indoor scenes.
- Luminance Dominates Geometry Formation (https://arxiv.org/pdf/2610.00749): Investigates luminance vs. chroma. Evaluated on Mip-NeRF 360, Tanks and Temples, Deep Blending, Shiny Blender.
- GS-PQM (https://arxiv.org/pdf/2610.00195): A parameter-domain quality metric for compressed GS. Evaluated on the GScomp-QA dataset.
- DSSR-3D (https://arxiv.org/pdf/2610.00040): Decoupled reasoning for view-dependent referring segmentation. Introduces ViewRef-GS benchmark.
- EffGS (https://github.com/CypressLi01/EffGS): Acceleration framework with frequency-aware importance scoring. Evaluated on Mip-NeRF 360, Deep Blending, Tanks and Temples, Mill-19, UrbanScene3D, GauU-Scene.
- Lens Flare Removal and Reconstruction (https://arxiv.org/pdf/2609.39527): Fine-tunes a latent diffusion model and introduces a GS-based flare representation. Uses Flare7K++ and VFX benchmark.
- TSGL (https://arxiv.org/pdf/2609.38635): Teacher-Student Graph Learning for 3DGS compression. Evaluated on Mip-NeRF 360, Tanks and Temples, Deep Blending.
- StereoGaussians (https://arxiv.org/pdf/2609.38592): Predicts 3DGS from stereo pairs. Introduces SceneSplat-Stereo dataset and evaluates on StereoNVS, ReplicaGS.
- Beyond Monoscopic Viewing (https://arxiv.org/pdf/2609.38525): Studies 3DGS quality in VR. Proposes Union initialization (SfM + VGGT) and controlled user studies.
- Gaussian Stippling (https://arxiv.org/pdf/2609.38488): Sorting-free rendering with hybrid sampling. Evaluated on Mip-NeRF 360, Tanks and Temples, Deep Blending, DL3DV.
- PneuTac (https://arxiv.org/pdf/2609.38418): Unified MPM-Gaussian Splatting Simulation for soft robots. Uses Genesis physics simulator and Digit tactile sensor.
- OIC-GS (https://arxiv.org/pdf/2609.34367): Omnidirectional GS compression with HEALPix grids. Uses Flickr-sourced 360° dataset and SUN360.
- Imagine3D-LLM (https://cvlab-kaist.github.io/Imagine3D-LLM): Multimodal LLM that imagines 3D scenes. Uses Gaussian summary tokens and distillation from ZipSplat.
- RLX (https://github.com/MIT-RLX/rlx-paper): Unified multi-backend tensor compiler in Rust. Supports 14+ runtime devices and 100+ model-family crates.
- EndoPrior-GS (https://jiaqi-huang-77.github.io/EndoPrior-GS/): Dynamic endoscopic reconstruction with a joint texture prior. Evaluated on EndoNeRF and SCARED benchmarks.
- WINGS (https://arxiv.org/pdf/2609.37816): Reference-free 3DGS inpainting using a pretrained 3D generative prior (TRELLIS).
- NRF-GS (https://arxiv.org/pdf/2609.37115): Neural Residual Fields for expressive and compact GS. Evaluated on Mip-NeRF 360, DL3DV, Tanks and Temples, BTF.
- DispFlow-GS (https://arxiv.org/pdf/2609.36940): Displacement Flow Supervision for deformable 3DGS. Introduces Deformation-Rendering Consistency (DRC) metric. Evaluated on NeRF-DS, HyperNeRF.
- AESplat (https://github.com/aesplat/AESplat): Decoupled appearance modeling for pose-free feed-forward 3DGS. Evaluated on RealEstate10K, ACID, DL3DV, ScanNet++.
- Distilling Privileged CBF (https://syeon-yoo.github.io/distill-cbf-site/): RGB-only safety filters for visual navigation. Uses Gaussian Splatting for a privileged teacher.
- GenNVS (https://arxiv.org/pdf/2609.34579): Geometry-enhanced Novel View Synthesis via Disentangled 3D Prior. Uses SpatialVid-HQ for training and Mip-NeRF 360, DL3DV for evaluation.
Impact & The Road Ahead
These advancements signify a pivotal shift for Gaussian Splatting, transforming it from a powerful rendering technique into a cornerstone for a wide array of AI/ML applications. The ability to generate dynamic, looping 3D content (OuroWorld) opens new avenues for entertainment and creative industries. Enhancements in sparse-view reconstruction (Sparse-View 4D Gaussian Splatting, UGOD, DensiTok, StereoGaussians) will accelerate 3D capture in real-world settings with limited sensors.
The push for improved geometric fidelity and semantic understanding (PCAsplat, Prior-Driven Enhancements, MEGA, OpenSplatGraph, DSSR-3D) is crucial for robotics, augmented reality, and virtual reality, where accurate interactions and object-level reasoning are paramount. Imagine robots navigating and interacting with environments described not just visually, but semantically, powered by robust 3DGS maps. The development of specialized benchmarks like GaussianBench and ViewRef-GS is key to ensuring these systems are not just visually plausible but physically and semantically correct.
Efficiency gains (Budgeted-GS, ByteSplat, TileSkipper, EffGS, GSCV) mean that 3DGS is becoming increasingly viable for large-scale, real-time applications, from city-scale digital twins to mobile AR/VR. The ability to efficiently compress 3DGS data, whether through video codecs or graph-based methods (TSGL), is vital for widespread deployment.
Beyond direct technical improvements, tools like SPLATIFY are democratizing research, accelerating the conversion of ideas into reproducible code and fostering compositional innovation. The insights into how geometry is formed (Luminance Dominates Geometry Formation) and how to evaluate quality in immersive contexts (Beyond Monoscopic Viewing) are fundamental contributions to the underlying science.
The future of 3DGS appears bright, poised to enable more intelligent, interactive, and immersive experiences across industries. From giving MLLMs the capacity to “imagine 3D scenes” (Imagine3D-LLM) to allowing tactile manipulation with soft robots (PneuTac) and even enabling privacy-preserving multi-robot collaboration (TRACE), Gaussian Splatting is clearly becoming a foundational technology, pushing the boundaries of what’s possible in 3D AI.
Share this content:
Discover more from SciPapermill
Subscribe to get the latest posts sent to your email.
Post Comment