Structure matters
Point clouds become more compressible when their physical or acquisition structure is made visible to the codec.
I investigated preconditioning techniques for 3D point clouds to improve their compressibility—starting with image-based representations and progressing to structure-aware ZFP experiments.
Method 1 encodes coordinates into image channels, then lets a lossless image codec do the work. Stanford Bunny, 35,947 vertices.
435.289 kB → 388.656 kB · −10.71% · median XYZ error 0.0617
This clean reproduction uses the 35,947-vertex Stanford Bunny. The original float32 XYZ payload is reshaped into an image-compatible byte layout, encoded, decoded, and rebuilt as a point cloud. Rotate the result to see how lossy image compression can leave shadows or aliases of the geometry.
Interactive representation experiment
Rotate the Stanford Bunny, compare the original XYZ cloud against the PNG- and JPEG-decoded payloads, and overlay all three to inspect geometric shadows and aliases.
Two fingers to rotate · pinch to zoom
PNG reconstructed this payload bit-for-bit. JPEG quality 100 changed 42.4% of payload bytes; its median XYZ error was 0.0617 in the source coordinate units.
A point cloud is often stored as an unordered list of XYZ values. Generic compressors see neighbouring rows in memory, not neighbouring points in space. My work tests whether image layouts, ordering, spherical coordinates, compact indices, and controlled precision can expose enough structure for numerical compressors to become effective.
LiDAR produces large spatial datasets for autonomous driving, mapping, digital twins, and robotics. Compression can reduce storage and network demand, but file size alone is an incomplete metric. A fair experiment must also measure geometric fidelity, attribute recovery, ordering requirements, decoder completeness, and runtime.
Point clouds become more compressible when their physical or acquisition structure is made visible to the codec.
Each experiment isolates image layout, ordering, score, index coding, quantization, or numerical compression.
Payload, framing, permutations, precision, reconstruction error, and decoder requirements are evaluated together.
If XYZ bytes are reshaped into an image carrier, PNG or JPEG may exploit local patterns even without understanding geometry.
The binary XYZ payload is reshaped into an image-compatible byte layout. PNG is tested losslessly; JPEG quality 100 is tested as an explicitly lossy control.
Image wrapping reduced the stored payload by 10.71% with PNG and 3.93% with JPEG q100. The clean Bunny reproduction then exposes the reconstruction consequences.
What this shows. PNG found more byte-level redundancy than JPEG q100, while the interactive Bunny makes the geometric cost of the lossy path visible.
Sorting by radial distance or a structural score should make neighbouring numerical values easier for compression codecs to encode.
Original-order, distance-sorted, and score-sorted XYZ streams are compared. Sorted variants also store the permutation needed to restore point order.
Sorting improved one component of the encoded stream, but the permutation needed to restore point order cost more than the component saved. The complete order-recoverable archive was larger than the baseline.
Points ordered by a recovered log-cardinality score may reveal regions that a numerical codec encodes more efficiently.
All points, the lower-score half, and the higher-score half are evaluated along a controlled removal sequence with signed byte changes.
The score did separate easier from harder subsets, but neither subset compressed below its raw size. Informative as a diagnostic, not sufficient as a codec — and the one-by-one recompression cost grows roughly quadratically.
Quantized spherical angles repeat. Replacing them with indices should expose smaller, more regular values.
Phi and theta indices are repeatedly centred, converted to absolute deltas, and paired with explicit sign bits until four-bit magnitudes remain.
The iterative midpoint-delta representation with explicit sign packing produced the smallest self-contained encoding of the six exact variants tested, with exact index recovery.
Distances, angle dictionaries, index deltas, and signs have different statistical structure and can use different encoders.
XYZ is quantized in spherical form; repeated distances are run-length coded; angle indices use the midpoint-delta transform; numerical and byte streams are encoded separately.
Quantization supplied the larger first reduction; stream-specific custom coding added a second, smaller one. Separating the two contributions was the point of the experiment.
If a raw XYZ matrix already contains sufficient numerical correlation, reversible ZFP should reduce it without a representation change.
The raw float64 coordinate matrix is compressed directly using reversible ZFP and decoded exactly.
Applying reversible ZFP directly to the unordered coordinate matrix produced a stream larger than the original payload. A rectangular array in memory is not a smooth spatial field — which is exactly what motivates the representation study.
LASzip/LAZ and Draco provided domain-specific reference points against which the custom experiments were measured.
Mapping coordinate axes into image channels, beyond the single-carrier layout shown in Method 1.
Iterative reordering strategies aimed at bringing spatially related points into a more regular sequence before coding.
Local filtering and smoothing approaches trading geometric fidelity for compressibility.
Large 3D streams affect logging, mapping, remote robotics, simulation, and continuous model development. Current reference directions solve different parts of the problem rather than producing one universal winner.
Directly codes sparse 3D geometry and attributes. Its generality brings configuration, complexity, and rate–distortion choices.
Uses the video ecosystem for dynamic volumetric content, but projection and patch processing add artifacts and overhead.
Efficiently transports meshes and point clouds, while quantization and graphics-oriented assumptions must match the downstream task.
Exploit acquisition order and temporal range-image structure, which is powerful but sensor- and representation-dependent.
Uses spherical LiDAR structure with learned compression; model cost, deployment complexity, and domain transfer remain practical considerations.
Excels on structured numerical fields, but unordered particle lists do not naturally supply the smooth neighbourhoods it expects.
No single method simultaneously optimizes sparse and dense geometry, attributes, temporal prediction, exact order, random access, streaming latency, bounded error, and downstream perception quality. The central engineering task is to match representation, fidelity budget, and access pattern to the application—and measure the complete recoverable system.