Eyes on Many: Evaluating Gaze, Hand, and Voice for Multi-Object Selection in Extended Reality

1 Concordia University • 2 University of Calgary • 3 Aarhus University
Eyes on Many: Evaluating Gaze, Hand, and Voice for Multi-Object Selection in Extended Reality teaser

Overview of four mode-switching and three subselection techniques for gaze-based multi-object selection in XR.


Abstract

Interacting with multiple objects simultaneously makes us fast. A pre-step to this interaction is to select the objects, i.e., multi-object selection, which is enabled through two steps: (1) toggling multi-selection mode - mode-switching - and then (2) selecting all the intended objects - subselection. In extended reality (XR), each step can be performed with the eyes, hands, and voice. To examine how design choices affect user performance, we evaluated four mode-switching (SemiPinch, FullPinch, DoublePinch, and Voice) and three subselection techniques (Gaze+Dwell, Gaze+Pinch, and Gaze+Voice) in a user study. Results revealed that while DoublePinch paired with Gaze+Pinch yielded the highest overall performance, SemiPinch achieved the lowest performance. Although Voice-based mode-switching showed benefits, Gaze+Voice subselection was less favored, as the required repetitive vocal commands were perceived as tedious. Overall, these findings provide empirical insights and inform design recommendations for multi-selection techniques in XR.

Methodology

SemiPinch

Multi-selection is activated when the distance between thumb and index fingertips is within 2–7 cm (a partial pinch pose). Users maintain this grip to stay in multi-selection mode while subselecting targets. The grouping is finalized with a full-pinch. Releasing the semi-pinch reverts to single-selection mode.

FullPinch

The user maintains a full-pinch (thumb and index fingertips < 2 cm apart) to activate and sustain multi-selection mode. When the pinch is released (> 7 cm apart), the system waits 250 ms before deactivating the mode and finalizing the group.

DoublePinch

A double-pinch is recognized when a full-pinch is released and performed again within 350 ms. This persistently toggles multi-selection mode on. Another double-pinch deactivates the mode and finalizes the selected group.

Voice

Participants use a spoken command (e.g., "group", "multi") to activate multi-selection, and a distinct term (e.g., "done", "finish") to deactivate and finalize. This provides a hands-free persistent mode-switching option.

Results

DoublePinch paired with Gaze+Pinch yielded the highest overall performance in terms of task completion time, error rate, and inverse efficiency. SemiPinch produced the highest mode and selection error rate, lowest efficiency, and was rated as the most fatiguing. Voice-based mode-switching showed benefits for hands-free interaction, but Gaze+Voice subselection was less favored due to the tedium of repetitive vocal commands. 17 out of 30 participants preferred DoublePinch for mode-switching, and 15 preferred Gaze+Dwell for subselection.

BibTeX

@inproceedings{bashar2026eyes,
  author = {Bashar, Mohammad Raihanul and Mutasim, Aunnoy K. and Pfeuffer, Ken and Batmaz, Anil Ufuk},
  title = {Eyes on Many: Evaluating Gaze, Hand, and Voice for Multi-Object Selection in Extended Reality},
  booktitle = {Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems},
  series = {CHI '26},
  articleno = {671},
  numpages = {14},
  year = {2026},
  doi = {10.1145/3772318.3790513}
}