Research: iPhone Pro LiDAR capture to a metric, gravity-aligned splat #5

Closed
opened 2026-08-11 15:41:38 +00:00 by lars · 1 comment
Owner

Question

Which capture pipeline turns an iPhone Pro (LiDAR) walkthrough of an office room into a splat that is real-world scale and gravity-aligned, and in a format the chosen viewer loads?

Answer specifically:

  • Tools (Polycam, Luma, Scaniverse, Postshot, Nerfstudio, Reality Composer / Object Capture) — which produce gaussian splats rather than meshes, and which preserve metric scale from LiDAR rather than an arbitrary one.
  • Whether output is gravity-aligned, and how to align it if not.
  • Export formats and how they map onto the viewer options from the WKWebView research.
  • Cost / subscription and whether export is watermark-free and offline.
  • Practical capture technique for an indoor room: lighting, path, overlap, how long a capture takes, and typical failure modes.
  • How the splat's own coordinate frame relates to anything measurable in the room — the input to the alignment step.

Part of the wayfinder map #1.

## Question Which capture pipeline turns an iPhone Pro (LiDAR) walkthrough of an office room into a splat that is **real-world scale and gravity-aligned**, and in a format the chosen viewer loads? Answer specifically: - Tools (Polycam, Luma, Scaniverse, Postshot, Nerfstudio, Reality Composer / Object Capture) — which produce gaussian splats rather than meshes, and which preserve metric scale from LiDAR rather than an arbitrary one. - Whether output is gravity-aligned, and how to align it if not. - Export formats and how they map onto the viewer options from the WKWebView research. - Cost / subscription and whether export is watermark-free and offline. - Practical capture technique for an indoor room: lighting, path, overlap, how long a capture takes, and typical failure modes. - How the splat's own coordinate frame relates to anything measurable in the room — the input to the alignment step. --- Part of the wayfinder map #1.
lars added this to the Wayfinder: RN prototyping spec milestone 2026-08-11 15:41:38 +00:00
lars added the wayfinder:research label 2026-08-11 15:41:38 +00:00
lars added a new dependency 2026-08-11 15:42:12 +00:00
Author
Owner

Answer

Research date: 2026-08-11. Every claim below is tagged [VERIFIED] (read directly in a vendor doc / repo / source file), [VENDOR CLAIM] (stated by the vendor but unverifiable without a test), or [COMMUNITY] (forum/review/third-party). Nothing here has been tested on hardware yet - the taped-point test described at the end is what converts the vendor claims into verified ones.

1. Which tools actually produce gaussian splats, and which preserve metric scale

Tool 3DGS? Where processed Metric scale from ARKit/LiDAR? Gravity-aligned? Splat export Free / offline / watermark Cost (2026-08-11)
Scaniverse Classic (Niantic) Yes [VERIFIED] On-device [VERIFIED, App Store text] Yes - "metric scale support" listed as a free-tier feature [VENDOR CLAIM] Undocumented; inherits ARKit frame, so probably yes [UNVERIFIED] SPZ, PLY (+FBX/GLB/USDZ mesh) [VERIFIED] Free, offline, no watermark found Classic free/unlimited. New cloud Scaniverse: Free $0 (20k credits/mo, no commercial rights), Plus $20/mo-$200/yr, Pro $50/mo-$500/yr (commercial rights)
Polycam Yes [VERIFIED] Cloud No for splats. Splat (Object) mode "does not require a LiDAR sensor"; MEASUREMENTS are documented only for Space/Floorplan modes [VERIFIED]. Their own Unity guide says "You may need to adjust the scale and rotation" [VERIFIED] Manual only: global export setting "Splat up axis: Y up / Y down / Z up" [VERIFIED]. No auto-level Splat PLY only, gated to Business/Enterprise [VERIFIED] Cloud; free tier watermarks video/images Free $0; Basic $150/yr or $30/mo; Business $300/yr/user (needed for splat export); Enterprise $1,200/yr/seat, 3-seat min
Polycam raw Space-mode LiDAR (not the splat product) n/a - this is input data On-device export, Developer Mode Yes, metric. polyform docs: poses follow "the ARKit gravity aligned convention", depth PNGs are 16-bit millimetres [VERIFIED] Yes, explicitly - RH, +Y up, -Z = session-start view dir, perpendicular to gravity [VERIFIED] Feeds ns-process-data polycam -> transforms.json Free; images only retrievable on the original device Free
Luma AI 3D Capture Yes [VENDOR CLAIM] Cloud No - "No Lidar ... necessary" [VERIFIED] Undocumented PLY [COMMUNITY only] - App free, no IAP - but effectively abandoned: v1.3.14 last updated 2026-01-14 and its release note is a Genie sunset notice; changelog has no 3D-capture entry since 2024-11; lumalabs.ai homepage no longer surfaces 3D capture. Do not build on this.
KIRI Engine Yes [VERIFIED] Cloud only - local processing "isn't on our roadmap" [VERIFIED] No for splats. LiDAR Scan is a separate mode exporting only OBJ/USDZ [VERIFIED]. v4.2.0 (2026-04-10) added "measurement and rescaling" for mesh projects only [VERIFIED] - a rescale tool exists precisely because output isn't metric Undocumented; they publish an FAQ titled "Why Can't I Properly View 3DGS Files in 3D Editing Software?" [VERIFIED] PLY only Cloud Free tier (150 photos/2 GB); Pro $17.99/mo or $79.99/yr on site vs $9.99-17.99/mo, $29.99-49.99/yr on App Store IAP - discrepancy, verify in-app
Jawset Postshot Yes (Splat3, Splat MCMC) [VERIFIED] Local, offline Preserved if you feed known poses - with imported poses/points "Postshot will skip the image selection and camera tracking stages" i.e. no SfM, no rescale [VERIFIED]. No ARKit importer; you must convert to COLMAP/Bundler/RealityScan/BlocksExchange/Pix4D Manual: "Align Up Vector" + Transform + Bake Transform on the Rdnc Field node [VERIFIED]. No auto gravity PLY (with SH), SPZ v4, HTML export. Export gated to Indie tier Free tier = non-commercial + watermark + no field export Indie EUR 17/mo or EUR 204/yr; Studio EUR 39/mo or EUR 468/yr [COMMUNITY - jawset.com/shop/pricing rendered EUR 0.00 on fetch, JS-populated; check in a browser]. Windows-only, NVIDIA CC >= 7.5
Nerfstudio splatfacto + gsplat Yes [VERIFIED] Local Destroyed by default - see section 1a. Recoverable with flags orientation_method="up" by default (estimates up from camera vectors) [VERIFIED] splat.ply with f_dc_*/f_rest_* SH, opacity, scale_0-2, rot_0-3 [VERIFIED from exporter.py] Apache 2.0, offline, no watermark Free. CUDA NVIDIA GPU, ~8+ GB VRAM. ~8-20 min/room on a 4090 [COMMUNITY]
Brush (ArthurBrussee) Yes (MCMC) [VERIFIED] Local Preserved verbatim - source inspection of crates/brush-dataset/src/{config.rs,formats/nerfstudio.rs,scene.rs} shows no pose scaling, no centering, no auto-orient, and it ignores applied_transform [VERIFIED by source] None automatic. Manual in-viewer: grid widget + arrow keys to level the ground (0.3.0) Plain .ply with SH (export_{iter}.ply). No SPZ/SOG writer Apache 2.0, offline, no watermark Free. Cross-platform incl. Apple Silicon via wgpu - no CUDA. v0.3.0, 2025-09-14
Apple Object Capture / PhotogrammetrySession No - mesh only (USDZ/OBJ) [VERIFIED] On-device n/a n/a n/a Free Free
RealityKit GaussianSplatComponent (WWDC26) Render only, not a trainer [VERIFIED] n/a n/a n/a No file format - you supply raw buffers -> BufferResource -> GaussianSplatResource Free SDK Free

1a. The nerfstudio metric-scale trap (most important single finding)

NerfstudioDataParserConfig defaults [VERIFIED from source]: orientation_method="up", center_method="poses", auto_scale_poses=True, load_3D_points=False. The offending line is literally:

scale_factor /= float(torch.max(torch.abs(poses[:, :3, 3])))

i.e. the scene is normalised to fit +/-1. Worse: ns-export gaussian-splat does not apply the inverse dataparser transform (unlike ExportPointCloud), so there is no post-hoc rescue - you must train unnormalised:

ns-train splatfacto --data <dir> nerfstudio-data \
  --auto-scale-poses False --orientation-method none \
  --center-method none --scale-factor 1.0 --load-3D-points True

ns-process-data record3d applies no scaling and no unit conversion [VERIFIED from record3d_utils.py] - ARKit metres pass through into transforms.json unchanged. Nerfstudio's own docs state scale is "preserved from the LiDAR depth data" for record3d and "real-world scale maintained" for polycam raw [VERIFIED].

2. Gravity alignment - and how to fix it if absent

ARKit's world frame with ARWorldAlignment.gravity is already gravity-aligned: +Y opposite gravity, metres, X/Z arbitrary in the horizontal plane (.gravityAndHeading additionally pins -Z to north). So gravity alignment is free at capture time - the risk is a pipeline stage throwing it away, not the sensor.

Consequences:

  • Do not use nerfstudio's orientation_method="up". It estimates up from average camera vectors, which is strictly worse than the gravity vector ARKit already gave you. Use none and apply a fixed basis swap yourself.
  • The Y-up -> Z-up question is a fixed constant basis change for the whole pipeline, not a per-scene fit. Decide it once against the chosen viewer.
  • If a splat does arrive un-levelled, fixing it is a one-line offline op - splat-transform -r <x,y,z> (Euler degrees), -t <x,y,z>, -s <factor> [VERIFIED]. SuperSplat's browser editor can do it interactively. Postshot has "Align Up Vector"; Brush 0.3.0 has a grid+arrow-key leveller.

Note that rotating a splat is not just moving positions: the per-splat rotation quaternions and the spherical-harmonic coefficients must be rotated too. SPZ handles this explicitly via Wigner D-matrices on save/load [VERIFIED]. Use a real splat tool, never a naive point-cloud transform.

3. Export formats

Format What it is Notes
.ply (splat PLY, f_dc_*/f_rest_*/opacity/scale_0-2/rot_0-3) The de-facto lossless interchange master. Large (up to GB) Every trainer writes it; every tool reads it. Use this as the archival master.
.splat Simple quantised binary from antimatter15's viewer Widely readable, no SH beyond DC in the common variant
.ksplat mkkellogg/GaussianSplats3D's own format - "a trimmed-down and compressed version of the original .ply", matching the viewer's internal layout for "the fastest loading times" [VERIFIED] Effectively viewer-specific; converted via that repo's browser UI or util/create-ksplat.js
SOG / Streamed SOG PlayCanvas's web-delivery format, "15-20x smaller than PLY, lossy" [VERIFIED]; Streamed SOG loads on-demand chunks Best web-delivery choice if you land on the PlayCanvas engine. Every SuperSplat publish is compressed to SOG
SPZ Niantic's open format, MIT, "~10x smaller than the corresponding .ply" [VERIFIED]. Now v4 (ZSTD, independent attribute streams); v1-3 still readable Best-documented coordinate story of any format: default RUB = +Y up, right-handed, OpenGL/three.js convention, contrasted in-doc with PLY's RDF, GLB's LUF, Unity's RUF; 16 named coordinate systems, source/dest specifiable on save and load with SH rotation [VERIFIED]

Format choice does not need to block the viewer ticket. PlayCanvas's splat-transform (open source, npm CLI, runs offline) reads PLY, compressed PLY, SOG, Streamed SOG, SPZ, .splat, .ksplat, LCC/LCC2 and writes PLY, compressed PLY, SOG, Streamed SOG, SPZ, GLB, CSV, HTML viewer, LOD, voxel, WebP [VERIFIED]. So: keep a metric, gravity-correct .ply master and transcode to whatever the viewer ticket picks. The only lock-in risk is .ksplat, which points at one specific viewer.

4. Indoor capture technique for one office room

Postshot's Capturing Guidelines is the best primary source anyone publishes [all VERIFIED, quoted]:

  • Path: "two to five rings around the scene, each with a different tilt ... and pedestal" (i.e. multiple loops at different heights and tilts - knee, chest, overhead). For a room specifically: stand against the wall and shoot along the longest sightline into the space. Capture floor, ceiling and surroundings, not just the objects.
  • Never rotate in place: "Don't pan, tilt or roll without moving the camera at the same time" - triangulation needs baseline. This is the single most common indoor failure.
  • Overlap: "Aim for about 30-50% overlap between images." Frame count 40 (controlled studio) to "several hundreds or thousands" (complex scenes). Postshot typically samples only 2-3 FPS from video, so walking speed matters more than frame rate - use the lowest available framerate.
  • Sharpness beats noise: "Prefer short exposure times and small apertures to avoid both motion blur and defocus blur"; if forced to choose, favour the small aperture; "Radiance fields tend to tolerate noise better than blur" - so let ISO climb.
  • Lens: ~24 mm full-frame equivalent handheld; "most stable reconstruction ... from images that have a good amount of context rather than close-up shots."
  • Lighting must be strictly static: no moving illumination, no on-camera flash (differential illumination as the camera moves), minimise lens flares ("extreme form of view-dependent lighting").
  • Scene must be static: "Anything that moves on the images will create blur or ghosting artefacts." Nerfstudio adds: capture "overlapping, non-blurry images"; Polycam recommends good lighting and moving slowly, or manual shutter mode.

Failure modes to expect in an office room - note that no vendor documents these, so this is engineering judgement plus the physics of 3DGS, not a cited claim:

  • Windows - the worst offender. Bright exteriors blow out and defeat auto-exposure; LiDAR returns nothing through glass. Close blinds; if you can't, capture at dusk or accept a hole.
  • Monitors and screens - turn them off. A live screen is a moving light source.
  • Reflective/specular surfaces (whiteboards, glass tabletops, glossy laminate) - 3DGS bakes reflections as view-dependent SH; expect floaters and a "second room" behind the glass.
  • Blank walls - no texture means no SfM features. This is exactly where LiDAR-pose ingest (which does not depend on photometric features for pose) beats a COLMAP-only pipeline.
  • Moving people - clear the room. One walk-through leaves ghosts.
  • LiDAR range ceiling ~5 m, and Polycam's own docs warn sub-1-2 cm detail is unreliable [VERIFIED] - fine for a single office, a hard limit for a long corridor (relevant to the eventual ship application).
  • Loop closure / drift - end where you started and re-observe the first wall. Polycam's raw poses are globally optimised (bundle-adjusted) and nerfstudio calls them "more robust to drift (an issue with ARKit or SLAM methods)" [VERIFIED] - a real reason to prefer Polycam raw over raw VIO for anything larger than one room.

Practical recipe: ~4-8 minutes of walking, 3 loops, ~300-800 sampled frames.

5. Relating splat space to the physical room (the similarity transform)

The splat's own frame is whatever the trainer emitted. To bridge to a measured site frame you need a similarity transform p_site = s*R*p_splat + t - 7 DOF (3 rotation, 3 translation, 1 uniform scale) - which is the classic Umeyama / Procrustes least-squares problem (Umeyama 1991, IEEE TPAMI).

What that requires concretely:

  • >= 3 non-collinear correspondences to determine the transform; use 4-6 so you have residuals to inspect. Spread them across the room and across all three axes (don't put all markers on the floor - coplanar points make the fit ill-conditioned about the horizontal axes).
  • Each correspondence is a pair: a point measured in the room (tape measure / laser, in metres) and the same point identified in splat space (click it in the viewer). Sharp physical corners are far easier to click than tape crosses on flat floor - use table corners, door frame corners, a plumb-bob marker at each taped point.
  • Off-the-shelf solver: Open3D's TransformationEstimationPointToPoint(with_scaling=True), documented as producing T = [[cR, t], [0, 1]] where with_scaling=False "force[s] scaling to be 1" [VERIFIED]. So the same API gives you both the similarity and the rigid fit.

Does metric capture remove the scale term? Effectively yes, but keep fitting it anyway. If the capture is genuinely metric, s should come out at 1.000 +/- a small tolerance, and you should then lock s = 1 and fit only R, t (6 DOF, rigid) for production - fitting a free scale against noisy hand-clicked points will absorb real error into s and flatter your residuals. But on the first capture, fit s free and read it as a measurement: s = 1.03 means the pipeline has a 3% scale error, and that number is the single most valuable output of the whole test. A pipeline where s drifts between captures is disqualified for AR navigation, because pose error grows with distance from the origin.

Also: if gravity alignment holds, R reduces to a single yaw angle plus a fixed basis swap. Fit the full R first and check that its roll/pitch are near zero - that is your empirical test of the gravity claim, and it is a much sharper test than eyeballing the render.

The taped-point test this ticket needs: tape 4-6 points on the floor, measure the pairwise distances between them with a tape measure before capturing (ground truth independent of any transform), capture, then (a) measure the same distances in splat space and compare - that alone validates metric scale with zero transform fitting - and (b) fit the transform, stand at each taped point, and compare render vs reality. Report RMS residual in millimetres.


Verdict

Capture tool: Scaniverse Classic (iOS). Export format: .ply as the metric master, transcoded with splat-transform to whatever the viewer ticket picks.

Rationale: it is the only app in the set with a vendor statement about metric scale, and it is simultaneously free, on-device/offline, watermark-free, and exports both PLY and SPZ. Polycam and KIRI both fail metric scale by construction (their splat modes don't use LiDAR at all) and Polycam additionally paywalls splat export at $300/yr. Luma is abandoned. .ply as the master is the choice that keeps the viewer ticket genuinely open - splat-transform converts it to .splat, .ksplat, SOG, Streamed SOG, SPZ or compressed PLY offline, so no viewer decision is foreclosed. Avoid .ksplat as a master: it points at one specific viewer.

Fallback if Scaniverse's metric scale fails the taped-point test: Polycam Space-mode raw LiDAR export (Developer Mode) -> ns-process-data polycam -> Brush. This is the rigorous path: Polycam's polyform docs explicitly specify the ARKit gravity-aligned convention with millimetre depth, and Brush is verified by source inspection not to touch your poses (no rescale, no re-centre, no auto-orient) - unlike nerfstudio, whose defaults silently destroy metric scale. Brush is Apache 2.0 and runs on Apple Silicon without CUDA, so it needs no new hardware. Postshot would give better quality but costs EUR 17/mo and is Windows+NVIDIA only.

Capture recipe for the office room:

  1. Prep: clear all people. Turn every monitor and screen off. Close blinds - windows are the top failure mode. Turn on all room lights and leave them fixed; no flash, no moving lamps, no daylight changes mid-capture.
  2. Tape 4-6 ground-truth points spread across the room and at differing heights where possible. Put a clickable physical feature at each (small box corner or plumb-bob marker), not just a tape cross on flat floor. Measure all pairwise distances with a tape measure and write them down now - this is your transform-free scale check.
  3. Start against a wall, facing along the longest sightline into the room (Postshot's explicit room advice). Note the start pose - this is ARKit's origin, +Y up, metres.
  4. Loop 1 - chest height, walk the perimeter slowly, camera aimed slightly down and outward into the room. Keep translating continuously; never pan or tilt while standing still.
  5. Loop 2 - knee height, same perimeter, tilted up. Loop 3 - overhead, tilted down, to get the floor and desk tops. Three rings at different tilt and pedestal, per Postshot's "two to five rings" rule.
  6. Maintain 30-50% overlap; walk at a pace that keeps frames sharp. Prefer high ISO over any motion blur.
  7. Close the loop: finish where you started and re-observe the first wall for several seconds to bound drift.
  8. Add a short pass over each taped marker so every ground-truth point is densely observed from >= 3 directions - sparse coverage there directly inflates your residuals.
  9. Total ~4-8 min. Process on-device, export .ply (and SPZ as a second copy for comparison).
  10. Validate, in this order: (a) measure the taped pairwise distances in splat space -> confirms metric scale with no transform fitting; (b) fit Umeyama with with_scaling=True -> read s (expect ~1.000) and check R's roll/pitch ~ 0 -> confirms gravity alignment; (c) lock s=1, refit rigid, stand at each taped point and compare render vs reality; (d) report RMS residual in mm.

Open / unresolved

  • Scaniverse's "metric scale support" is an unverified vendor claim on a pricing page - no technical doc, no stated accuracy figure. The old scaniverse.com help pages (which described the ruler tool) now 302-redirect to nianticspatial.com, so those primaries are gone. The taped-point test is the only way to settle this, and the whole recommendation rests on it.
  • Does Scaniverse gravity-align / auto-level its splat output? No documentation either way, for any of the four consumer apps. Inferred from the ARKit frame only. A 2026-08-04 App Store review complains "It's been over 2 years and still no support to rotate splats" [COMMUNITY] - so if it arrives un-levelled there is no in-app fix and you must use splat-transform -r.
  • Scaniverse Classic vs new Scaniverse is a genuine contradiction in the vendor's own material. The App Store text says "unlimited on-device Gaussian splat capture and processing"; the nianticspatial.com FAQ says "Scans are uploaded, then processed in the cloud." These describe different products sharing a name; Classic accounts explicitly do not carry over ("You will need to create a new account"). Confirm which binary you actually have before trusting the offline/on-device property.
  • Commercial-rights licensing. New Scaniverse Free and Plus tiers are marked "Commercial rights: No"; commercial use starts at Pro $50/mo. Scaniverse Classic's terms are a separate document at scaniverse.com/terms and were not read. For a client-facing cruise-ship product this needs a real answer before the prototype becomes a deliverable.
  • Postshot pricing unconfirmed - jawset.com/shop/pricing rendered EUR 0.00 for all three tiers on fetch (JS-populated or a live promo). The EUR 17 / EUR 39 figures are third-party. Only matters if the fallback path escalates to Postshot.
  • No tool in the consumer set exports SOG or .ksplat, and only Scaniverse exports SPZ. All web-delivery formats therefore require an offline splat-transform step. Fine, but it means the pipeline is never a single app.
  • splat-transform's SH rotation correctness was not verified by source. SPZ documents Wigner D-matrix SH rotation explicitly; splat-transform's -r was only read from its README. If gravity correction is applied via -r, verify view-dependent appearance doesn't degrade.
  • Nothing here is hardware-tested. In particular: actual splat quality on blank office walls, whether LiDAR's ~5 m range is limiting in the specific test room, and how badly a single unavoidable window degrades the result.
  • Scale-stability across captures is untested and is the real gate for AR navigation. One good s ~ 1.000 proves little; the requirement is that s be reproducible across repeated captures of the same room, since pose error grows with distance from the origin. Plan >= 3 captures.
  • Corridor / full-ship scale is out of scope here and does not follow from a one-room result. LiDAR range, drift and loop closure all degrade with extent; Polycam's bundle-adjusted poses (nerfstudio: "more robust to drift") are the likely answer at that scale, which may mean the ship pipeline is not the office pipeline.

Primary sources (all fetched 2026-08-11): Polycam capture modes - Polycam export file types - Polycam splats to Unity - Polycam pricing - PolyCam/polyform - Scaniverse App Store - Niantic Spatial pricing - Scaniverse FAQ - nianticlabs/spz - KIRI export formats - KIRI local processing FAQ - KIRI pricing - Luma 3D Capture App Store - Luma changelog - Postshot Capturing Guidelines - Postshot Importing Images - Postshot Rdnc Field node - nerfstudio custom data - nerfstudio_dataparser.py - camera_utils.py - exporter.py - record3d_utils.py - gsplat LICENSE - ArthurBrussee/brush - PlayCanvas splat formats - playcanvas/splat-transform - SuperSplat - mkkellogg/GaussianSplats3D - Open3D TransformationEstimationPointToPoint - RealityKit GaussianSplatComponent - Stray Scanner format - jc211/NeRFCapture

## Answer **Research date: 2026-08-11.** Every claim below is tagged **[VERIFIED]** (read directly in a vendor doc / repo / source file), **[VENDOR CLAIM]** (stated by the vendor but unverifiable without a test), or **[COMMUNITY]** (forum/review/third-party). Nothing here has been tested on hardware yet - the taped-point test described at the end is what converts the vendor claims into verified ones. ### 1. Which tools actually produce gaussian splats, and which preserve metric scale | Tool | 3DGS? | Where processed | Metric scale from ARKit/LiDAR? | Gravity-aligned? | Splat export | Free / offline / watermark | Cost (2026-08-11) | |---|---|---|---|---|---|---|---| | **Scaniverse Classic** (Niantic) | Yes [VERIFIED] | **On-device** [VERIFIED, App Store text] | **Yes** - "metric scale support" listed as a free-tier feature [VENDOR CLAIM] | Undocumented; inherits ARKit frame, so probably yes [UNVERIFIED] | **SPZ, PLY** (+FBX/GLB/USDZ mesh) [VERIFIED] | Free, offline, no watermark found | Classic free/unlimited. New cloud Scaniverse: Free $0 (20k credits/mo, **no commercial rights**), Plus $20/mo-$200/yr, Pro $50/mo-$500/yr (commercial rights) | | **Polycam** | Yes [VERIFIED] | Cloud | **No for splats.** Splat (Object) mode "does not require a LiDAR sensor"; MEASUREMENTS are documented only for Space/Floorplan modes [VERIFIED]. Their own Unity guide says "You may need to adjust the **scale and rotation**" [VERIFIED] | Manual only: global export setting "Splat up axis: Y up / Y down / Z up" [VERIFIED]. No auto-level | Splat **PLY only**, gated to **Business/Enterprise** [VERIFIED] | Cloud; free tier watermarks video/images | Free $0; Basic $150/yr or $30/mo; **Business $300/yr/user** (needed for splat export); Enterprise $1,200/yr/seat, 3-seat min | | **Polycam raw Space-mode LiDAR** (not the splat product) | n/a - this is *input data* | On-device export, Developer Mode | **Yes, metric.** `polyform` docs: poses follow "the **ARKit gravity aligned convention**", depth PNGs are 16-bit **millimetres** [VERIFIED] | **Yes, explicitly** - RH, +Y up, -Z = session-start view dir, perpendicular to gravity [VERIFIED] | Feeds `ns-process-data polycam` -> `transforms.json` | Free; images only retrievable on the original device | Free | | **Luma AI 3D Capture** | Yes [VENDOR CLAIM] | Cloud | **No** - "No Lidar ... necessary" [VERIFIED] | Undocumented | PLY [COMMUNITY only] | - | App free, no IAP - but **effectively abandoned**: v1.3.14 last updated 2026-01-14 and its release note is a *Genie sunset* notice; changelog has no 3D-capture entry since 2024-11; lumalabs.ai homepage no longer surfaces 3D capture. **Do not build on this.** | | **KIRI Engine** | Yes [VERIFIED] | **Cloud only** - local processing "isn't on our roadmap" [VERIFIED] | **No for splats.** LiDAR Scan is a *separate mode* exporting only OBJ/USDZ [VERIFIED]. v4.2.0 (2026-04-10) added "measurement and **rescaling**" for **mesh** projects only [VERIFIED] - a rescale tool exists precisely because output isn't metric | Undocumented; they publish an FAQ titled "Why Can't I Properly View 3DGS Files in 3D Editing Software?" [VERIFIED] | **PLY only** | Cloud | Free tier (150 photos/2 GB); Pro $17.99/mo or $79.99/yr on site vs $9.99-17.99/mo, $29.99-49.99/yr on App Store IAP - **discrepancy, verify in-app** | | **Jawset Postshot** | Yes (Splat3, Splat MCMC) [VERIFIED] | Local, offline | **Preserved if you feed known poses** - with imported poses/points "Postshot will skip the image selection and camera tracking stages" i.e. no SfM, no rescale [VERIFIED]. No ARKit importer; you must convert to COLMAP/Bundler/RealityScan/BlocksExchange/Pix4D | Manual: "Align Up Vector" + Transform + Bake Transform on the Rdnc Field node [VERIFIED]. No auto gravity | PLY (with SH), **SPZ v4**, HTML export. Export gated to Indie tier | Free tier = non-commercial + **watermark** + no field export | Indie EUR 17/mo or EUR 204/yr; Studio EUR 39/mo or EUR 468/yr [COMMUNITY - jawset.com/shop/pricing rendered EUR 0.00 on fetch, JS-populated; check in a browser]. **Windows-only, NVIDIA CC >= 7.5** | | **Nerfstudio splatfacto + gsplat** | Yes [VERIFIED] | Local | **Destroyed by default** - see section 1a. Recoverable with flags | `orientation_method="up"` by **default** (estimates up from camera vectors) [VERIFIED] | `splat.ply` with `f_dc_*`/`f_rest_*` SH, opacity, `scale_0-2`, `rot_0-3` [VERIFIED from `exporter.py`] | **Apache 2.0**, offline, no watermark | Free. CUDA NVIDIA GPU, ~8+ GB VRAM. ~8-20 min/room on a 4090 [COMMUNITY] | | **Brush** (ArthurBrussee) | Yes (MCMC) [VERIFIED] | Local | **Preserved verbatim** - source inspection of `crates/brush-dataset/src/{config.rs,formats/nerfstudio.rs,scene.rs}` shows no pose scaling, no centering, no auto-orient, and it ignores `applied_transform` [VERIFIED by source] | None automatic. Manual in-viewer: grid widget + arrow keys to level the ground (0.3.0) | Plain `.ply` with SH (`export_{iter}.ply`). No SPZ/SOG writer | **Apache 2.0**, offline, no watermark | Free. **Cross-platform incl. Apple Silicon via wgpu - no CUDA.** v0.3.0, 2025-09-14 | | **Apple Object Capture / PhotogrammetrySession** | **No - mesh only** (USDZ/OBJ) [VERIFIED] | On-device | n/a | n/a | n/a | Free | Free | | **RealityKit `GaussianSplatComponent`** (WWDC26) | **Render only, not a trainer** [VERIFIED] | n/a | n/a | n/a | No file format - you supply raw buffers -> `BufferResource` -> `GaussianSplatResource` | Free SDK | Free | #### 1a. The nerfstudio metric-scale trap (most important single finding) `NerfstudioDataParserConfig` defaults [VERIFIED from source]: `orientation_method="up"`, `center_method="poses"`, **`auto_scale_poses=True`**, `load_3D_points=False`. The offending line is literally: ```python scale_factor /= float(torch.max(torch.abs(poses[:, :3, 3]))) ``` i.e. the scene is normalised to fit +/-1. Worse: **`ns-export gaussian-splat` does not apply the inverse dataparser transform** (unlike `ExportPointCloud`), so there is no post-hoc rescue - you must train unnormalised: ``` ns-train splatfacto --data <dir> nerfstudio-data \ --auto-scale-poses False --orientation-method none \ --center-method none --scale-factor 1.0 --load-3D-points True ``` `ns-process-data record3d` applies **no scaling and no unit conversion** [VERIFIED from `record3d_utils.py`] - ARKit metres pass through into `transforms.json` unchanged. Nerfstudio's own docs state scale is "preserved from the LiDAR depth data" for record3d and "real-world scale maintained" for polycam raw [VERIFIED]. ### 2. Gravity alignment - and how to fix it if absent ARKit's world frame with `ARWorldAlignment.gravity` is already gravity-aligned: **+Y opposite gravity, metres**, X/Z arbitrary in the horizontal plane (`.gravityAndHeading` additionally pins -Z to north). So gravity alignment is *free at capture time* - the risk is a pipeline stage throwing it away, not the sensor. Consequences: - **Do not use nerfstudio's `orientation_method="up"`.** It *estimates* up from average camera vectors, which is strictly worse than the gravity vector ARKit already gave you. Use `none` and apply a fixed basis swap yourself. - The Y-up -> Z-up question is a **fixed constant basis change** for the whole pipeline, not a per-scene fit. Decide it once against the chosen viewer. - If a splat does arrive un-levelled, fixing it is a one-line offline op - `splat-transform -r <x,y,z>` (Euler degrees), `-t <x,y,z>`, `-s <factor>` [VERIFIED]. SuperSplat's browser editor can do it interactively. Postshot has "Align Up Vector"; Brush 0.3.0 has a grid+arrow-key leveller. Note that rotating a splat is not just moving positions: the per-splat rotation quaternions **and the spherical-harmonic coefficients** must be rotated too. SPZ handles this explicitly via Wigner D-matrices on save/load [VERIFIED]. Use a real splat tool, never a naive point-cloud transform. ### 3. Export formats | Format | What it is | Notes | |---|---|---| | **`.ply`** (splat PLY, `f_dc_*`/`f_rest_*`/`opacity`/`scale_0-2`/`rot_0-3`) | The de-facto lossless interchange master. Large (up to GB) | Every trainer writes it; every tool reads it. **Use this as the archival master.** | | **`.splat`** | Simple quantised binary from antimatter15's viewer | Widely readable, no SH beyond DC in the common variant | | **`.ksplat`** | mkkellogg/GaussianSplats3D's own format - "a trimmed-down and compressed version of the original `.ply`", matching the viewer's internal layout for "the fastest loading times" [VERIFIED] | Effectively viewer-specific; converted via that repo's browser UI or `util/create-ksplat.js` | | **SOG / Streamed SOG** | PlayCanvas's web-delivery format, "**15-20x smaller than PLY**, lossy" [VERIFIED]; Streamed SOG loads on-demand chunks | Best web-delivery choice if you land on the PlayCanvas engine. Every SuperSplat publish is compressed to SOG | | **SPZ** | Niantic's open format, **MIT**, "~**10x smaller** than the corresponding `.ply`" [VERIFIED]. Now **v4** (ZSTD, independent attribute streams); v1-3 still readable | Best-documented coordinate story of any format: default **RUB = +Y up, right-handed, OpenGL/three.js convention**, contrasted in-doc with PLY's RDF, GLB's LUF, Unity's RUF; 16 named coordinate systems, source/dest specifiable on save and load with SH rotation [VERIFIED] | **Format choice does not need to block the viewer ticket.** PlayCanvas's `splat-transform` (open source, npm CLI, runs offline) reads PLY, compressed PLY, SOG, Streamed SOG, SPZ, **`.splat`, `.ksplat`**, LCC/LCC2 and writes PLY, compressed PLY, SOG, Streamed SOG, SPZ, GLB, CSV, HTML viewer, LOD, voxel, WebP [VERIFIED]. So: **keep a metric, gravity-correct `.ply` master and transcode to whatever the viewer ticket picks.** The only lock-in risk is `.ksplat`, which points at one specific viewer. ### 4. Indoor capture technique for one office room Postshot's *Capturing Guidelines* is the best primary source anyone publishes [all VERIFIED, quoted]: - **Path:** "two to five rings around the scene, each with a different tilt ... and pedestal" (i.e. multiple loops at different heights and tilts - knee, chest, overhead). For a room specifically: **stand against the wall and shoot along the longest sightline into the space.** Capture floor, ceiling and surroundings, not just the objects. - **Never rotate in place:** "Don't pan, tilt or roll without moving the camera at the same time" - triangulation needs baseline. This is the single most common indoor failure. - **Overlap:** "Aim for about **30-50% overlap** between images." Frame count 40 (controlled studio) to "several hundreds or thousands" (complex scenes). Postshot typically samples only **2-3 FPS** from video, so walking speed matters more than frame rate - use the lowest available framerate. - **Sharpness beats noise:** "Prefer short exposure times and small apertures to avoid both motion blur and defocus blur"; if forced to choose, favour the small aperture; "Radiance fields tend to **tolerate noise better than blur**" - so let ISO climb. - **Lens:** ~24 mm full-frame equivalent handheld; "most stable reconstruction ... from images that have a good amount of **context** rather than close-up shots." - **Lighting must be strictly static:** no moving illumination, **no on-camera flash** (differential illumination as the camera moves), minimise lens flares ("extreme form of view-dependent lighting"). - **Scene must be static:** "Anything that moves on the images will create blur or ghosting artefacts." Nerfstudio adds: capture "overlapping, non-blurry images"; Polycam recommends good lighting and moving slowly, or manual shutter mode. **Failure modes to expect in an office room** - note that no vendor documents these, so this is engineering judgement plus the physics of 3DGS, not a cited claim: - **Windows** - the worst offender. Bright exteriors blow out and defeat auto-exposure; LiDAR returns nothing through glass. Close blinds; if you can't, capture at dusk or accept a hole. - **Monitors and screens** - turn them **off**. A live screen is a moving light source. - **Reflective/specular surfaces** (whiteboards, glass tabletops, glossy laminate) - 3DGS bakes reflections as view-dependent SH; expect floaters and a "second room" behind the glass. - **Blank walls** - no texture means no SfM features. This is exactly where LiDAR-pose ingest (which does not depend on photometric features for pose) beats a COLMAP-only pipeline. - **Moving people** - clear the room. One walk-through leaves ghosts. - **LiDAR range ceiling ~5 m**, and Polycam's own docs warn sub-1-2 cm detail is unreliable [VERIFIED] - fine for a single office, a hard limit for a long corridor (relevant to the eventual ship application). - **Loop closure / drift** - end where you started and re-observe the first wall. Polycam's raw poses are **globally optimised (bundle-adjusted)** and nerfstudio calls them "more robust to drift (an issue with ARKit or SLAM methods)" [VERIFIED] - a real reason to prefer Polycam raw over raw VIO for anything larger than one room. **Practical recipe: ~4-8 minutes of walking, 3 loops, ~300-800 sampled frames.** ### 5. Relating splat space to the physical room (the similarity transform) The splat's own frame is whatever the trainer emitted. To bridge to a measured site frame you need a **similarity transform** `p_site = s*R*p_splat + t` - 7 DOF (3 rotation, 3 translation, 1 uniform scale) - which is the classic **Umeyama / Procrustes** least-squares problem (Umeyama 1991, *IEEE TPAMI*). What that requires concretely: - **>= 3 non-collinear correspondences** to determine the transform; use **4-6** so you have residuals to inspect. Spread them across the room and across all three axes (don't put all markers on the floor - coplanar points make the fit ill-conditioned about the horizontal axes). - Each correspondence is a pair: a point measured in the room (tape measure / laser, in metres) and the *same* point identified in splat space (click it in the viewer). Sharp physical corners are far easier to click than tape crosses on flat floor - use table corners, door frame corners, a plumb-bob marker at each taped point. - Off-the-shelf solver: Open3D's `TransformationEstimationPointToPoint(with_scaling=True)`, documented as producing `T = [[cR, t], [0, 1]]` where `with_scaling=False` "force[s] scaling to be 1" [VERIFIED]. So the same API gives you both the similarity and the rigid fit. **Does metric capture remove the scale term? Effectively yes, but keep fitting it anyway.** If the capture is genuinely metric, `s` should come out at 1.000 +/- a small tolerance, and you should then *lock* `s = 1` and fit only `R, t` (6 DOF, rigid) for production - fitting a free scale against noisy hand-clicked points will absorb real error into `s` and flatter your residuals. But on the **first** capture, fit `s` free and read it as a **measurement**: `s = 1.03` means the pipeline has a 3% scale error, and that number is the single most valuable output of the whole test. A pipeline where `s` drifts between captures is disqualified for AR navigation, because pose error grows with distance from the origin. Also: if gravity alignment holds, `R` reduces to a **single yaw angle** plus a fixed basis swap. Fit the full `R` first and check that its roll/pitch are near zero - that is your empirical test of the gravity claim, and it is a much sharper test than eyeballing the render. **The taped-point test this ticket needs:** tape 4-6 points on the floor, measure the pairwise distances between them with a tape measure *before* capturing (ground truth independent of any transform), capture, then (a) measure the same distances in splat space and compare - that alone validates metric scale with **zero** transform fitting - and (b) fit the transform, stand at each taped point, and compare render vs reality. Report RMS residual in millimetres. --- ### Verdict **Capture tool: Scaniverse Classic (iOS). Export format: `.ply` as the metric master, transcoded with `splat-transform` to whatever the viewer ticket picks.** Rationale: it is the **only** app in the set with a vendor statement about **metric scale**, and it is simultaneously **free, on-device/offline, watermark-free, and exports both PLY and SPZ**. Polycam and KIRI both fail metric scale *by construction* (their splat modes don't use LiDAR at all) and Polycam additionally paywalls splat export at $300/yr. Luma is abandoned. `.ply` as the master is the choice that keeps the viewer ticket genuinely open - `splat-transform` converts it to `.splat`, `.ksplat`, SOG, Streamed SOG, SPZ or compressed PLY offline, so no viewer decision is foreclosed. Avoid `.ksplat` as a master: it points at one specific viewer. **Fallback if Scaniverse's metric scale fails the taped-point test:** Polycam Space-mode **raw** LiDAR export (Developer Mode) -> `ns-process-data polycam` -> **Brush**. This is the rigorous path: Polycam's `polyform` docs *explicitly* specify the ARKit gravity-aligned convention with millimetre depth, and Brush is verified **by source inspection** not to touch your poses (no rescale, no re-centre, no auto-orient) - unlike nerfstudio, whose defaults silently destroy metric scale. Brush is Apache 2.0 and runs on Apple Silicon without CUDA, so it needs no new hardware. Postshot would give better quality but costs EUR 17/mo and is Windows+NVIDIA only. **Capture recipe for the office room:** 1. **Prep:** clear all people. Turn every monitor and screen **off**. Close blinds - windows are the top failure mode. Turn on all room lights and leave them fixed; no flash, no moving lamps, no daylight changes mid-capture. 2. **Tape 4-6 ground-truth points** spread across the room and at differing heights where possible. Put a clickable physical feature at each (small box corner or plumb-bob marker), not just a tape cross on flat floor. **Measure all pairwise distances with a tape measure and write them down now** - this is your transform-free scale check. 3. **Start against a wall, facing along the longest sightline** into the room (Postshot's explicit room advice). Note the start pose - this is ARKit's origin, +Y up, metres. 4. **Loop 1 - chest height,** walk the perimeter slowly, camera aimed slightly down and outward into the room. Keep translating continuously; **never pan or tilt while standing still.** 5. **Loop 2 - knee height,** same perimeter, tilted up. **Loop 3 - overhead,** tilted down, to get the floor and desk tops. Three rings at different tilt and pedestal, per Postshot's "two to five rings" rule. 6. **Maintain 30-50% overlap;** walk at a pace that keeps frames sharp. Prefer high ISO over any motion blur. 7. **Close the loop:** finish where you started and re-observe the first wall for several seconds to bound drift. 8. **Add a short pass over each taped marker** so every ground-truth point is densely observed from >= 3 directions - sparse coverage there directly inflates your residuals. 9. Total ~4-8 min. Process on-device, **export `.ply`** (and SPZ as a second copy for comparison). 10. **Validate, in this order:** (a) measure the taped pairwise distances in splat space -> confirms metric scale with no transform fitting; (b) fit Umeyama with `with_scaling=True` -> read `s` (expect ~1.000) and check `R`'s roll/pitch ~ 0 -> confirms gravity alignment; (c) lock `s=1`, refit rigid, stand at each taped point and compare render vs reality; (d) report RMS residual in mm. ### Open / unresolved - **Scaniverse's "metric scale support" is an unverified vendor claim on a pricing page** - no technical doc, no stated accuracy figure. The old scaniverse.com help pages (which described the ruler tool) now 302-redirect to nianticspatial.com, so those primaries are **gone**. The taped-point test is the only way to settle this, and the whole recommendation rests on it. - **Does Scaniverse gravity-align / auto-level its splat output?** No documentation either way, for any of the four consumer apps. Inferred from the ARKit frame only. A 2026-08-04 App Store review complains "It's been over 2 years and still no support to rotate splats" [COMMUNITY] - so if it arrives un-levelled there is no in-app fix and you must use `splat-transform -r`. - **Scaniverse Classic vs new Scaniverse is a genuine contradiction in the vendor's own material.** The App Store text says "unlimited **on-device** Gaussian splat capture and processing"; the nianticspatial.com FAQ says "Scans are **uploaded, then processed in the cloud**." These describe different products sharing a name; Classic accounts explicitly do **not** carry over ("You will need to create a new account"). **Confirm which binary you actually have** before trusting the offline/on-device property. - **Commercial-rights licensing.** New Scaniverse Free *and* Plus tiers are marked "Commercial rights: No"; commercial use starts at Pro $50/mo. Scaniverse **Classic's** terms are a separate document at scaniverse.com/terms and were **not** read. For a client-facing cruise-ship product this needs a real answer before the prototype becomes a deliverable. - **Postshot pricing unconfirmed** - jawset.com/shop/pricing rendered EUR 0.00 for all three tiers on fetch (JS-populated or a live promo). The EUR 17 / EUR 39 figures are third-party. Only matters if the fallback path escalates to Postshot. - **No tool in the consumer set exports SOG or `.ksplat`**, and only Scaniverse exports SPZ. All web-delivery formats therefore require an offline `splat-transform` step. Fine, but it means the pipeline is never a single app. - **`splat-transform`'s SH rotation correctness was not verified by source.** SPZ documents Wigner D-matrix SH rotation explicitly; splat-transform's `-r` was only read from its README. If gravity correction is applied via `-r`, verify view-dependent appearance doesn't degrade. - **Nothing here is hardware-tested.** In particular: actual splat quality on blank office walls, whether LiDAR's ~5 m range is limiting in the specific test room, and how badly a single unavoidable window degrades the result. - **Scale-stability across captures is untested and is the real gate for AR navigation.** One good `s ~ 1.000` proves little; the requirement is that `s` be reproducible across repeated captures of the same room, since pose error grows with distance from the origin. Plan >= 3 captures. - **Corridor / full-ship scale is out of scope here and does not follow from a one-room result.** LiDAR range, drift and loop closure all degrade with extent; Polycam's bundle-adjusted poses (nerfstudio: "more robust to drift") are the likely answer at that scale, which may mean the ship pipeline is *not* the office pipeline. --- **Primary sources** (all fetched 2026-08-11): [Polycam capture modes](https://learn.poly.cam/hc/en-us/articles/48565771018772-Which-Capture-Mode-Should-I-Use) - [Polycam export file types](https://learn.poly.cam/hc/en-us/articles/27756102599572-What-File-Types-Can-Polycam-Export) - [Polycam splats to Unity](https://learn.poly.cam/hc/en-us/articles/41491673295508-How-to-Import-Your-Gaussian-Splat-Captures-into-Unity) - [Polycam pricing](https://poly.cam/pricing) - [PolyCam/polyform](https://github.com/PolyCam/polyform) - [Scaniverse App Store](https://apps.apple.com/us/app/scaniverse-3d-scanner/id1541433223) - [Niantic Spatial pricing](https://www.nianticspatial.com/pricing) - [Scaniverse FAQ](https://www.nianticspatial.com/faq/scaniverse) - [nianticlabs/spz](https://github.com/nianticlabs/spz) - [KIRI export formats](https://www.kiriengine.app/features/export-formats) - [KIRI local processing FAQ](https://www.kiriengine.app/faq/run-algorithm-locally) - [KIRI pricing](https://www.kiriengine.app/pricing) - [Luma 3D Capture App Store](https://apps.apple.com/us/app/luma-3d-capture/id1615849914) - [Luma changelog](https://lumalabs.ai/changelog) - [Postshot Capturing Guidelines](https://www.jawset.com/docs/d/Postshot+User+Guide/Capturing+Guidelines) - [Postshot Importing Images](https://www.jawset.com/docs/d/Postshot+User+Guide/Importing+Images) - [Postshot Rdnc Field node](https://www.jawset.com/docs/d/Postshot+User+Guide/Interface/Scene+Tree/Rdnc+Field) - [nerfstudio custom data](https://docs.nerf.studio/quickstart/custom_dataset.html) - [nerfstudio_dataparser.py](https://github.com/nerfstudio-project/nerfstudio/blob/main/nerfstudio/data/dataparsers/nerfstudio_dataparser.py) - [camera_utils.py](https://github.com/nerfstudio-project/nerfstudio/blob/main/nerfstudio/cameras/camera_utils.py) - [exporter.py](https://github.com/nerfstudio-project/nerfstudio/blob/main/nerfstudio/scripts/exporter.py) - [record3d_utils.py](https://github.com/nerfstudio-project/nerfstudio/blob/main/nerfstudio/process_data/record3d_utils.py) - [gsplat LICENSE](https://github.com/nerfstudio-project/gsplat/blob/main/LICENSE) - [ArthurBrussee/brush](https://github.com/ArthurBrussee/brush) - [PlayCanvas splat formats](https://developer.playcanvas.com/user-manual/gaussian-splatting/formats/) - [playcanvas/splat-transform](https://github.com/playcanvas/splat-transform) - [SuperSplat](https://developer.playcanvas.com/user-manual/supersplat/) - [mkkellogg/GaussianSplats3D](https://github.com/mkkellogg/GaussianSplats3D) - [Open3D TransformationEstimationPointToPoint](https://www.open3d.org/docs/release/python_api/open3d.pipelines.registration.TransformationEstimationPointToPoint.html) - [RealityKit GaussianSplatComponent](https://developer.apple.com/documentation/realitykit/gaussiansplatcomponent) - [Stray Scanner format](https://docs.strayrobots.io/apps/scanner/format.html) - [jc211/NeRFCapture](https://github.com/jc211/NeRFCapture)
lars closed this issue 2026-08-11 15:59:53 +00:00
Sign in to join this conversation.