Skip to content

Vision scanning and voxel carving

Learn how silhouettes, visual cones and voxel carving help ASCAND estimate an object’s occupied volume from structured views.

Vision scanning and voxel carving

ASCAND can use a sequence of object silhouettes to estimate which parts of a three-dimensional volume may be occupied.

The principle is called voxel carving. Processing begins with a bounded volume around the object and divides that space into small three-dimensional cells called voxels. Each orientation-indexed image contributes a silhouette: a separation between the object and the visible background. Voxels that contradict those observations can be removed. Voxels that remain consistent across the available views survive.

The resulting occupied volume approximates the object’s visual hull: the largest three-dimensional form consistent with all observed silhouettes.

This is a powerful source of shape evidence, but it is not a complete digital copy. Silhouettes describe outer constraints. They do not directly measure surface depth, reveal internal structure or recover a concavity that never changes the object’s outline.

See how structured rotation supplies orientation context

A silhouette says where material may—and may not—exist

A silhouette divides an image into two broad regions:

  • Foreground: pixels classified as belonging to the object.
  • Background: pixels classified as lying outside the object.

Consider a viewing ray extending from the camera through one pixel and into the reconstruction region. If the pixel is reliably classified as background, object material should not occupy the corresponding ray through that region. A voxel projected into that background area therefore contradicts the observation and can be removed.

A foreground pixel has a different meaning. It indicates that the camera saw the object somewhere along that line of sight. It does not identify the exact surface depth by itself. Many voxels at different depths can project into the same foreground pixel and remain possible after one view.

That is why one silhouette creates a cone-like region of possible occupancy rather than a complete 3D surface. Views from additional orientations introduce additional constraints. Their intersection progressively narrows the remaining volume.

Silhouette evidence must also be kept separate from other information in the image:

  • An outline constrains occupied space.
  • Surface texture can support feature correspondence.
  • Pixel color can contribute appearance information.
  • Shading may suggest shape, but it is affected by lighting and material behavior.
  • None of these observations exposes hidden internal structure automatically.

The quality of foreground/background separation therefore matters directly. If part of the object is mistaken for background, valid volume may be carved away. If background is mistaken for the object, impossible volume may survive.

Assess object suitability

From structured views to occupied volume

A public conceptual model of the voxel-carving path contains six stages.

1. Prepare orientation-indexed views

The source video provides an ordered sequence of observations. The controlled camera–turntable relationship and visible coded reference help associate usable views with rotational context.

2. Separate the object from the background

Processing identifies the object region in each usable view. The result is a set of silhouettes or foreground masks. Clear visual separation supports this stage; changing shadows, reflections, transparency, motion and a similar object/background appearance can make it harder.

3. Initialize a bounded voxel grid

A three-dimensional region large enough to contain the object is divided into voxels. At the start, each voxel represents possible occupancy rather than confirmed material.

4. Project voxels into each silhouette

The acquisition geometry relates voxels in reconstruction space to positions in each image. Processing asks whether a candidate voxel is consistent with the observed object region from that orientation.

5. Retain, weaken or remove occupancy

Background evidence contradicts occupancy. Supporting or uncertain observations are accumulated according to the reconstruction logic. No isolated frame should be treated as a complete description of the object.

6. Preserve the volumetric estimate

After the available views have contributed their constraints, the surviving occupied region forms a voxel-derived estimate of the object. Later stages may assign appearance, extract a boundary surface, evaluate quality or compare this result with independent evidence.

This sequence explains the principle, not the current interface or its proprietary implementation. Exact thresholds, grid settings and processing decisions belong to verified technical documentation rather than a conceptual article.

Review the documented processing paths

Carving removes contradictions instead of collecting surface points

Voxel carving differs from reconstruction methods that accumulate surface samples.

Feature-based photogrammetric processing looks for relationships among visible image features and uses their multi-view geometry to infer surface structure. Laser triangulation observes the position of a projected laser line and derives surface samples from known geometric relationships. Voxel carving begins with possible occupied space and removes regions that cannot agree with the silhouettes.

Evidence methodPrimary observationReconstruction operation
Feature-based visionCorresponding image features across viewsInfer spatial surface relationships
Laser triangulationObserved laser line under known geometryDerive surface samples
Silhouette-based voxel carvingForeground and background across viewsEliminate impossible volume

These methods answer different geometric questions. Their outputs should not be treated as interchangeable merely because they originate from the same capture sequence.

The cumulative nature of carving is important. One view may leave a large range of depths possible. Another view removes some of those possibilities. Further orientations continue testing the remaining voxels. Regions supported by diverse observations have a different evidence basis from regions constrained by only a small number of similar views.

The process can also propagate errors. An inaccurate silhouette can remove material that really exists or preserve space that should be empty. Incorrect acquisition geometry can project voxels into the wrong image regions. Object movement relative to the turntable can make otherwise clear silhouettes inconsistent with the assumed model.

Repeatability of the algorithm does not make the input evidence infallible.

How ASCAND relates to photogrammetry

Voxels are an intermediate volume, not necessarily the final file

A voxel is a small region of three-dimensional space. Voxel size controls how finely the reconstruction volume can represent spatial change.

A coarser grid uses fewer cells. It can preserve broad form with lower computational and memory requirements, but small features may be simplified, merged or lost. A finer grid can represent smaller changes in the occupied boundary, while requiring more computation and memory.

Higher voxel resolution does not create missing evidence. If a recess never changes any observed silhouette, subdividing the volume more finely cannot reveal that recess. If a thin feature was removed by segmentation, a finer grid cannot restore it from the absent observation.

The occupied grid is also not identical to every downstream representation:

  1. Occupied volume: the voxels that remain plausible.
  2. Volumetric boundary: the transition between occupied and empty cells.
  3. Extracted surface: polygonal geometry derived from that boundary.
  4. Refined mesh: a surface that may undergo method-specific cleanup or preparation.
  5. Output file: a representation encoded for a particular downstream tool.

Each transition can affect topology, smoothness, detail and apparent closure. A visually smooth mesh may be convenient, yet smoothing can conceal the block-like discretization of the source volume. A closed surface can support some downstream workflows, yet closure alone does not prove that every surface is correctly located.

Understand point clouds and meshes

What silhouette-based reconstruction handles well—and what it cannot infer

Voxel carving has useful characteristics precisely because it depends on different evidence from feature matching or laser triangulation.

Silhouette evidence can supportSilhouette evidence cannot establish by itself
Overall outer volume consistent with multiple viewsExact depth of every visible surface
Reconstruction that does not require distinctive surface textureDeep concavities that never affect an outline
Stability where patterns are repetitiveHidden undersides or occluded regions
Globally connected volumetric reasoningInternal cavities or internal structure
A naturally bounded or closed estimateCertified dimensions or metrology
Independent evidence for later comparisonFitness for printing, CAD or another use

The visual-hull constraint is easiest to understand with a bowl-shaped or indented object. If the indentation is visible only as shading inside an otherwise unchanged outline, its interior shape does not constrain the silhouettes. The visual hull can bridge across that region because a filled form remains consistent with every observed outline.

Thin structures create another challenge. A feature approaching the spatial scale of the voxel grid or the effective image segmentation may become unstable, thicken, join a neighboring feature or disappear. Transparent and reflective boundaries can make foreground separation ambiguous. Articulated or flexible parts can change during rotation and violate the assumption that every frame observes the same rigid form.

Coverage matters too. A turn around one vertical axis produces useful side views from the selected camera elevation. It does not automatically observe the underside, the top of every recess or surfaces hidden behind other geometry. A second approved object orientation may contribute complementary evidence, but separate reconstructions still require evaluation and a defined merge workflow.

Review ASCAND’s capability boundaries

What the method requires from the object and capture

The reconstruction principle leads to practical capture consequences.

  • Object–background separation: The visible boundary should be distinguishable across the rotation. A good outline in one frame is not enough if it becomes ambiguous elsewhere.
  • Stable object placement: The object should not slide, flex or change pose relative to the turntable during the observation sequence.
  • Stable camera relationship: Unintended camera movement changes the acquisition geometry assumed across frames.
  • Controlled illumination: Flicker, moving shadows, glare and changing reflections can interfere with consistent foreground interpretation.
  • Useful focus and framing: The object and required reference regions must remain observable rather than being clipped, blurred or obscured.
  • Sufficient rotational evidence: The sequence must provide enough usable views to constrain the volume around the object.
  • Deliberate treatment of difficult features: Thin, transparent, reflective, dark, articulated or deeply recessed regions should be evaluated rather than assumed suitable.

These are conceptual dependencies, not a substitute for the current setup and Vision Scan tutorials. The verified procedure should remain the authority for interface actions, supported capture conditions and current operational steps.

The supplied turntable photograph can be used to confirm the product’s physical appearance and visible coded band. A complete capture illustration must additionally show the stationary smartphone in its real mounting relationship. The photograph alone does not demonstrate framing, phone placement, object visibility or a successful silhouette.

Review scan-video requirements
Set up the capture station

Choose the next evidence path

Voxel carving is best understood as one reconstruction path within a multi-evidence architecture. It contributes an occupied-volume estimate derived from silhouettes. Image correspondence, laser triangulation and complementary scans contribute different observations and different limitations.

Choose the next page according to your question:

The right question is not whether one method is universally best. It is which observable evidence constrains the geometry that matters for the intended result—and which uncertainty remains after processing.

Compare Vision, Laser and ComboHow ASCAND interprets structured rotation