🌲 SelectAnyTree

A Promptable Instance Segmentation Model for
3D Forest LiDAR Point Clouds

1Nagoya University 2RIKEN 3University of Freiburg 4UCLA 5University of Twente 6KC Machine Learning Lab 7HUST 8Ritsumeikan University
SelectAnyTree teaser — scene overview with three selected trees, top-down and side view

SelectAnyTree segments any individual tree in a 3D forest LiDAR scan with a user's click. Left: full scene overview with three highlighted trees. Right: zoom panels showing predicted masks. The scene is encoded once; each subsequent click is a cheap query.

Abstract

Instance segmentation of trees in forest LiDAR point clouds is constrained by label scarcity: a single hectare holds millions of points and hundreds of overlapping tree crowns, making manual annotation laborious, while automatic pre-segmentations offer no interactive refinement.

Inspired by the promptable paradigm of foundation segmentation models, we propose SelectAnyTree, a promptable instance segmentation model that delineates any individual tree in a 3D forest point cloud from a few clicks. It couples three lightweight stages: (1) a sparse voxel scene encoder that embeds the forest once into reusable features; (2) a click-to-query prompt encoder that turns each click into a single content query from its 3D position, positive/negative polarity, and the backbone feature of its nearest voxel; and (3) a state-space query decoder that converts this query into one tree mask in linear time, with a mask feedback that conditions each refinement round on the previous mask. Each additional tree costs only a lightweight prompt-encoding and decoding pass, and the full model uses just 19.4M parameters — far fewer than prior promptable 3D models. We further exploit forest-aware geometry by detecting treetops as local maxima of the Canopy Height Model (CHM) and associating one with the user's click as a free initial prompt.

Across seven diverse forest regions and an independent held-out dataset, SelectAnyTree segments a target tree to 79.9 IoU from a single click — 24.7 points above the strongest promptable baseline — and reaches every accuracy target with the fewest clicks.

Demo

Pre-computed from the test set. Each frame shows top-down view (left) and side view (right).

★ Positive click ✕ Negative click ▲ CHM treetop True positive False positive False negative

360° view — select any tree, click by click

Left: the whole plot with every tree segmented. Right: each tree, rotating, as prompts accumulate from 1 to 5 clicks.

SelectAnyTree — 360° rotating view of the scene and each segmented tree as clicks accumulate

Norway (NIBIO)

SelectAnyTree — Norway NIBIO, click-by-click segmentation

Australia (BlueCat)

SelectAnyTree — Australia BlueCat, click-by-click segmentation

Czech Republic (CULS)

SelectAnyTree — Czech Republic CULS, click-by-click segmentation

New Zealand (SCION)

SelectAnyTree — New Zealand SCION, click-by-click segmentation

Point-SAM vs. SelectAnyTree

Same scene, same click budget.

Point-SAM

Point-SAM comparison

SelectAnyTree Ours

SelectAnyTree comparison

Method

SelectAnyTree turns a structure-aware forest backbone into a promptable instance segmentation model through four stages:

  1. Sparse voxel scene encoder (once). The point cloud is voxelized and embedded by a sparse voxel encoder into per-voxel features — cached and reused for all subsequent clicks.
  2. Click-to-query prompt encoder. Each click set is mapped to a single content query from the backbone feature of the click's nearest voxel, a random-Fourier positional encoding of its coordinate, and signed aggregation of positive/negative click polarities.
  3. State-space query decoder. The query is decoded into a voxel-resolution mask by a state-space (Mamba) decoder that captures long-range context in linear time; a lightweight mask feedback conditions each refinement round on the previous mask.
  4. CHM-guided first prompt. Canopy Height Model treetops, detected from scene geometry, provide a "free" first prompt, bridging fully-automatic and interactive segmentation.
SelectAnyTree architecture

SelectAnyTree architecture. Scene encoding is performed once per plot; prompt encoding and decoding are repeated per click set.

Results

79.9 IoU @ 1 click
+24.7 pts above best baseline
#1 fewest clicks at every IoU target
19.4M params — 16× smaller than Point-SAM

Interactive Segmentation

FOR-instanceV2 test set (in-distribution) and LAUTx (cross-dataset generalization, held-out). Bold = best, underline = second best.

Interactive segmentation comparison: IoU@Clicks and NoC@IoU on FOR-instanceV2 and LAUTx

Model Size & Inference Efficiency

GPU wall-clock time (ms) per scene for a k-click session on the FOR-instanceV2 test set.

Model size and inference efficiency comparison

Qualitative Results

Single-click segmentation compared against promptable baselines on the same target tree.

Single-click segmentation qualitative comparison against baselines

BibTeX

@article{nguyen2026selectanytree,
  title   = {SelectAnyTree: A Promptable Instance Segmentation Model for 3D Forest {LiDAR} Point Clouds},
  author  = {Nguyen, Trung Thanh and Lusk, Daniel and Gerberding, Kilian and Vajna-Jehle, Janusch and Vu, Tuan-Anh and Le, Duc Viet and Vo, Tu and Nguyen, Phi Le and Kawanishi, Yasutomo and Komamizu, Takahiro and Ide, Ichiro and Frey, Julian and Kattenborn, Teja},
  journal = {arXiv preprint arXiv:2606.27491},
  year    = {2026}
}