Open Access
Issue
Acta Acust.
Volume 10, 2026
Article Number 58
Number of page(s) 22
Section Virtual Acoustics
DOI https://doi.org/10.1051/aacus/2026056
Published online 07 July 2026

© The Author(s), Published by EDP Sciences, 2026

Licence Creative CommonsThis is an Open Access article distributed under the terms of the Creative Commons Attribution License (https://creativecommons.org/licenses/by/4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.

1 Introduction

Accurate auditory localization is fundamental to perceptual immersion in loudspeaker-based virtual acoustics. Combined crosstalk cancellation (CTC) and binaural reproduction systems can enhance spatial realism by eliminating the influence of undesired speaker-to-ear transmission paths and precisely shaping the binaural cues at the listener’s ears. Principally, the interaural time difference (ITD) and interaural level difference (ILD) can be controlled so that the perceived image corresponds to a desired direction. Because such systems are calibrated for a nominal head pose, small translations and rotations perturb the loudspeaker-to-ear transfer paths and thus the realized binaural cues, degrading spatial fidelity. To address this, some studies try to assess the optimal loudspeaker topologies that offer naturally greater sweet spot [15], while others utilize linear arrays to offer more degrees of freedom or employ head-tracking systems [6, 7]. Understanding how frontal, lateral, and rotational head movements affect ITD, IACC and ILD across typical loudspeaker geometries and acoustic conditions is of practical importance.

Speaker-based binaural reproduction is commonly pursued in two complementary ways. One approach aims to recreate the responses of target head-related transfer functions (HRTFs) at the eardrums [811]. A second analytically convenient approach imposes prescribed broadband ITD–ILD pairs on the ear signals to emulate the dominant localization cues [1116]. These pairs can either impose the true ITD–ILD information as extracted by prerecorded HRTFs depending on the virtual source position, or impose artificial ITD–ILD combinations that can still yield strong localization [11, 12]. Interaural information has also been incorporated in sound-field-control methods [17].

CTC operates by inverting a multichannel acoustic plant between loudspeakers and ears. Under head motion, changes in path lengths alter this plant, so that the equalization designed at the initial pose no longer exactly realizes the target ear signals. Most studies on CTC performance degradation under head motion have largely focused on channel separation (CS) [1, 2]. However, it has been reported that CS does not reliably predict the localization performance [1820]. ITD is a preferable metric to assess localization performance under head motion [35, 2123] as it is more directly connected to azimuthal perception, particularly in the low-mid frequency range. Thus, this study aims to specifically assess how head movements affect the measured ITD as well as interaural cross-correlation (IACC) under different loudspeaker topologies. High IACC is important for the perception of a unified sound image [24] and for accurate ITD discrimination [25]. While not the main objective of this study, the effect of head movements on the resulting ILD is also reported.

Optimal source distribution (OSD) and related analyses suggest a frequency-dependent trade-off considering the optimal loudspeaker span. Specifically, wider spans tend to improve low-frequency reproduction, whereas narrower spans can be advantageous at higher frequencies [26, 27]. Given that ITD is the dominant cue for horizontal localization over much of the low-mid frequency range (roughly 200–1400 Hz), this perspective would imply that larger-span loudspeaker geometries may be beneficial for preserving ITD-based spatial perception under practical perturbations such as head motion.

At the same time, much of the CTC literature that evaluates robustness under listener misalignment emphasizes channel separation (CS) as a primary performance measure, often favoring relatively small-span arrangements [1, 3] for all displacement types. Since CS has been shown to correlate weakly with localization performance as shown by Majdak et al. [18], it remains unclear whether the loudspeaker spans that are optimal for CS are also optimal for localization robustness. This motivates a systematic evaluation of loudspeaker topology from the perspective of ITD (and closely related interaural coherence), rather than relying on CS.

In this study, it is shown why CS alone is generally insufficient to predict localization-related robustness under head motion. It is demonstrated that even perfect CS can coexist with inaccurate ITD/IACC/ILD reproduction. Moreover, CS-based rankings can change substantially across spectral regions and may conflict with cue fidelity in the ITD-dominant band, motivating the combined reporting of ITD and IACC for a more perceptually relevant assessment.

Furthermore, the present work extends prior ITD-based robustness analyses which are typically limited to small-span loudspeaker arrangements and emphasize lateral misalignment [3, 4], by evaluating a broader range of loudspeaker spans and displacement types (frontal, lateral, and yaw rotation). By including large-span configurations that are rarely considered in the existing ITD/IACC literature, important qualitative changes in yaw-rotation behavior are revealed, including substantially improved ITD stability for large spans.

Additionally, a large range of target ITD–ILD combinations (representing different intended source directions) are evaluated to determine regimes over which topology-dependent robustness trends remain valid. Our results suggest that prior temporal evaluations using idealized targets (e.g., equally timed delta functions; [3]) do not capture the full range of performance changes.

The analysis is interpreted for both head-locked and world-locked reproduction scenarios as topology-dependent ITD variability under displacements may be perceptually advantageous or detrimental depending on whether the auditory scene is intended to remain fixed to the head or to the external world.

In addition to the above, the influence of secondary factors (speaker distance, head size, and rotation center) on ITD and IACC stability is assessed, and reflections in typical listening-room geometries are examined to determine the extent to which free-field results persist under reflective conditions. Prior robustness studies on head misalignment rarely consider the role of reflections, particularly when evaluating interaural-cue fidelity.

2 Human sensitivity to ITD deviations and perceptual implications

Humans are remarkably sensitive to ITD changes [28]. The smallest detectable ITD is on the order of 10 μs under optimal conditions [29]. Sensitivity is highest for frequencies below 1.4 kHz [21, 22]. At higher frequencies, ITD discrimination worsens drastically as phase information becomes ambiguous [30].

It is important to characterize the perceptual consequences of ITD robustness. For binaural reproduction, two spatial reference modes are commonly distinguished: (a) world-locked and (b) head-locked [31]. In world-locked reproduction, the objective is to render a perceptually stationary source such that, under head rotations or translations, the auditory image remains fixed in external space; this mode is well suited to cinematic and virtual-reality (VR) scenarios. By contrast, head-locked reproduction anchors the auditory image to the listener, causing it to rotate with changes in head orientation. This is advantageous in VR when a sound should follow the listener (e.g., a sound object that is held or worn by the listener).

For world-locked reproduction (a source fixed in external space), head motion should induce the corresponding changes in ITD/ILD so that the percept remains stationary in the room. By contrast, in head-locked reproduction, head motion should not change the rendered ITD/ILD – keeping the ear signals effectively constant – so the virtual image remains fixed relative to the head.

It is important to differentiate between these two scenarios to design the binaural system accordingly. Depending on the reproduction scenario (head-locked or world-locked), different loudspeaker configurations offer specific advantages. For quick reference the reader can directly refer to Section 7.2.

3 Effect of head movements on ITD

Considering a binaural set of signals that arrive at the two ears, preliminary examinations can be made regarding how (a) the relative position between the loudspeakers and the head, and (b) the type of head movements affect the ITD. The ITD can be defined as the time delay τ (s) corresponding to the integer sample lag k (i.e., with sampling rate f s , τ = k/f s ) that maximizes the normalized cross-correlation between the right- and left-ear signals [32]:

R x R , x L ( k ) = n x R ( n ) x L ( n + k ) n x R 2 ( n ) n x L 2 ( n ) · Mathematical equation: $$ \begin{aligned} R_{x_R,x_L}(k) = \frac{\sum _{n} x_R(n)\,x_L(n+k)}{\sqrt{\sum _{n} x_R^2(n)}\, \sqrt{\sum _{n} x_L^2(n)}}\cdot \end{aligned} $$(1)

The search range is typically limited to −2000 μs <  τ <  2000 μs, providing sufficient headroom for the most perceptually significant ITDs (typically ranging between [ − 800, 800] μs). The ITD (in samples) is expressed as

ITD = arg max k R x R , x L ( k ) . Mathematical equation: $$ \text{ITD} = \arg\max_k R_{x_R, x_L}(k). $$(2)

The peak value of RxR, xL(k) under the search window corresponds to the interaural cross-correlation (IACC). Note that with this formulation, when the signal at the right ear arrives sooner than the signal at the left, the occurring ITD is positive (convention used in this study). Considering a 2 × 2 multiple input – multiple output (MIMO) configuration (two loudspeakers and two microphones) under free-field conditions, the signals arriving at the left and right ears can be written as

x L ( n ) = d L ( n ) + cross L ( n ) , x R ( n ) = d R ( n ) + cross R ( n ) , Mathematical equation: $$ \begin{aligned} \begin{aligned} x_L(n)&= d_L(n) + {{\,\mathrm{cross}\,}}_L(n),\\ x_R(n)&= d_R(n) + {{\,\mathrm{cross}\,}}_R(n), \end{aligned} \end{aligned} $$(3)

where d R (n), d L (n) are the direct signals from the right and left loudspeakers to the respective ears, while cross R (n), cross L (n) are the crosstalk signals at each ear originating from the opposite loudspeaker (Fig. 1).

Thumbnail: Figure 1. Refer to the following caption and surrounding text. Figure 1.

Direct and crosstalk signals arriving in both ears under a binaural reproduction scenario.

Under binaural reproduction, the ear signals x R (n) and x L (n) realize a specified ITD and ILD target. For direct ITD–ILD control in binaural reproduction, the target ear signals can be defined as

X L ( ω ) = α ( ω ) X R ( ω ) , α ( ω ) = g ( ω ) e j ϕ ( ω ) , Mathematical equation: $$ \begin{aligned} X_L(\omega )=\alpha (\omega )X_R(\omega ),\qquad \alpha (\omega )=g(\omega )e^{-j\phi (\omega )}, \end{aligned} $$(4)

where g(ω) encodes a frequency-dependent ILD and ϕ(ω) encodes the interaural phase difference.

In this work we assume broadband (frequency-independent) cues over the analysis band. Thus, constant ILD corresponds to g(ω)≡g, while frequency-independent ITD corresponds to the linear-phase case ϕ(ω)=ωτ, where τ is a constant interaural time delay. Thus

α ( ω ) = g e j ω τ , Mathematical equation: $$ \begin{aligned} \alpha (\omega )=g\,e^{-j\omega \tau }, \end{aligned} $$(5)

where |α(ω)| = g and ∠α(ω)= − ωτ.

Even though static binaural information is sub-optimal compared to using HRTFs, this is a common alternative [16, 26, 33] and allows for a better evaluation of the ITD deviation under various head movements and speaker geometries types.

To evaluate how a head movement and loudspeaker geometry affect the measured ITD after a head displacement, we can utilize the interaural cross-correlation expression between x R (n) and x L (n). Substituting equations (3)–(1) gives

R x R , x L ( k ) = 1 a n = [ d R ( n ) d L ( n + k ) + d R ( n ) cross L ( n + k ) + cross R ( n ) d L ( n + k ) + cross R ( n ) cross L ( n + k ) ] , Mathematical equation: $$ \begin{aligned} R_{x_R,x_L}(k)&= \frac{1}{a}\sum _{n=-\infty }^{\infty }\Big [d_R(n)\,d_L(n+k) \nonumber \\&\quad + d_R(n)\,{{\,\mathrm{cross}\,}}_{L}(n+k) \nonumber \\&\quad + {{\,\mathrm{cross}\,}}_{R}(n)\,d_L(n+k) \nonumber \\&\quad + {{\,\mathrm{cross}\,}}_{R}(n)\,{{\,\mathrm{cross}\,}}_{L}(n+k) \Big ], \end{aligned} $$(6)

where a = n x R 2 ( n ) n x L 2 ( n ) Mathematical equation: $ a=\sqrt{\sum_n x_R^2(n)}\sqrt{\sum_n x_L^2(n)} $ is the normalization factor. Equation (6) can be conveniently expressed as:

R x R , x L ( k ) = 1 a ( R d R , d L ( k ) + R d R , cross L ( k ) + R cross R , d L ( k ) + R cross R , cross L ( k ) ) . Mathematical equation: $$ \begin{aligned} R_{x_{R},x_L}(k)&= \frac{1}{a}(R_{d_R,d_L}(k) + R_{d_R,{{\,\mathrm{cross}\,}}_{L}}(k) \nonumber \\&\quad + R_{{{\,\mathrm{cross}\,}}_{R},d_L}(k) + R_{{{\,\mathrm{cross}\,}}_{R},{{\,\mathrm{cross}\,}}_{L}}(k)). \end{aligned} $$(7)

This expansion is useful for analyzing how head movements impact the ITD sensitivity depending on loudspeaker geometry. Due to natural head shadowing, R d R , d L (k) is typically the dominant term, especially under a large loudspeaker span geometry where speakers are positioned on the left and right sides of the head.

Figures 2 and 3 illustrate simplified diagrams of large span and small span speaker configurations. Assuming a circular head with diametrically opposite ears, in the large span geometry of Figure 2a describing a rotational displacement, both direct paths dR and dL are equally delayed, and the crosstalk paths are equally shortened. Hence RdR, dL(k) and RcrossR, crossL(k) are not shifted by head rotation. The mixed terms RdR, crossL(k) and RcrossR, dL(k) are naturally weak at the initial head pose due to natural head shadowing, but their contribution grows with increased rotation. That is, after a certain degree of rotation it is impossible to retain binaural differences at the two ears with normal loudspeakers. Factors like speakers’ diameter, head width, speakers’ distance, and others, also affect the maximum allowable rotation level.

Thumbnail: Figure 2. Refer to the following caption and surrounding text. Figure 2.

Direct and cross path changes under (a) rotational and (b) lateral displacements for the large span 180° 2 × 2 listener configuration. The black dots correspond to the initial ear positions before a head displacement, and the orange dots to the displaced ear positions. Thick lines (stronger energy contribution) correspond to the direct paths, and thin lines to the crosstalk paths.

Thumbnail: Figure 3. Refer to the following caption and surrounding text. Figure 3.

Direct and cross path changes under (a) rotational and (b) lateral displacements for a 2 × 2 small span listener configuration.

By contrast, in the small span geometry of Figure 3a even small head rotations modify R d R , d L (k) appreciably: the arrival time of d R shortens while that of d L is delayed, and similar changes occur for the other terms of R x R , x L (k). This produces a substantial shift of the cross-correlation maximum ( Appendix A) and thus larger deviations in the measured ITD. Consequently, the topology in Figure 2a is expected to better retain the target ITD (as set at the initial head position) during rotations.

In the large span geometry under lateral displacements (Fig. 2b), direct and crosstalk paths change asymmetrically. For a left shift, d L arrives earlier while d R is delayed, moving the maximum of R d R , d L (k). Meanwhile, R cross R , cross L (k) shifts in the opposite direction, but because cross paths carry less energy (head shadowing), its influence on the total R x R , x L (k) is weaker and the direct-path shift dominates. In the small span geometry for the same lateral movement (Fig. 3b), both direct paths d R , d L acquire nearly the same additional delay, and the wavefront motion is approximately orthogonal to the head movement, so arrival-time changes in all terms are small. Thus R d R , d L (k) remains relatively stable, and the large span geometry is expected to show higher ITD sensitivity to lateral shifts than the small span geometry.

Even though these initial assessments provide an initial estimation of how the ITD is affected by the head movement type and loudspeaker configuration, they do not consider various other effects like head diffraction and room reflections. Furthermore, the excitation type and the drive effort can deteriorate the ITD robustness depending on the reproduction condition.

These observations challenge previous research findings that considered smaller span configurations robust against rotational displacements. For example, Parodi and Rubak [1] assessed topology robustness using CS-based criteria (i.e., absolute and relative sweet spot size). For rotational displacements, they reported very large sweet spots across all examined loudspeaker spans (12°–60°) for both criteria, while also the smallest span produced the largest relative sweet spot. This seems to indicate that CS may not be a suitable metric to assess localization performance as discussed in Introduction and specifically demonstrated in Section 4.

Takeuchi et al. [3] also performed ITD-IACC evaluations, but compared only two relatively small spans (10° and 60°) and concluded that the smallest span (10°) arrangement is slightly more robust to rotational misalignment compared to the largest span (60°). However, their span range did not include large-span topologies (e.g., 180° span – i.e., the large span 2 × 2 configuration discussed above), where the analysis predicts qualitatively different ITD-IACC robustness behavior. In particular, extrapolating their small-span trend would suggest that increasing span further should continue to degrade robustness, whereas the provided theoretical explanation shows that this does not hold once large-span configurations are included.

4 Effect of head movements on channel separation (CS)

A frequently used way to quantify the robustness of crosstalk-cancellation (CTC) systems under listener misalignment is channel separation (CS), which measures how strongly undesired inter-channel leakage is suppressed. This metric directly reflects residual crosstalk energy at the ears, and it is used to define relative and absolute sweet spots [1, 2] around the nominal head pose.

Let H(ω) denote the loudspeaker–ear transfer matrix under an initial head pose. This matrix is inverted by the CTC system [34] so that the influence of the acoustic paths between the loudspeakers and the ears, the reproduction and recording subsystems, and the room-imposed reflections is reduced [35], allowing direct control of the signals at the listener’s ears. Also, let C(ω) denote the inverse filters designed at this initial head position.

4.1 Definition of channel separation

The mapping from the desired binaural target signals Y(ω)=[Y L (ω) Y R (ω)] to the reproduced ear signals W(ω)=[W L (ω) W R (ω)] can be written as

W ( ω ) = R ( ω ) Y ( ω ) , R ( ω ) = H ( ω ) C ( ω ) = [ R 11 ( ω ) R 12 ( ω ) R 21 ( ω ) R 22 ( ω ) ] . Mathematical equation: $$ \begin{aligned} \mathbf W (\omega )&= \mathbf R (\omega )\mathbf Y (\omega ),\nonumber \\ \mathbf R (\omega )&= \mathbf H (\omega )\mathbf C (\omega )= \begin{bmatrix} R_{11}(\omega )&\quad R_{12}(\omega )\\ R_{21}(\omega )&\quad R_{22}(\omega ) \end{bmatrix}. \end{aligned} $$(8)

Ideally, R(ω)≈I at the design pose, where I is the 2 × 2 identity matrix. Under head motion, H(ω) changes and thus R(ω) departs from the identity. A standard definition of CS is [2]

CS L ( ω ) = 20 log 10 | R 12 ( ω ) | | R 11 ( ω ) | , CS R ( ω ) = 20 log 10 | R 21 ( ω ) | | R 22 ( ω ) | , Mathematical equation: $$ \begin{aligned} \mathrm{CS} _L(\omega )&= 20\log _{10}\frac{|R_{12}(\omega )|}{|R_{11}(\omega )|},\nonumber \\ \mathrm{CS} _R(\omega )&= 20\log _{10}\frac{|R_{21}(\omega )|}{|R_{22}(\omega )|}, \end{aligned} $$(9)

with more negative values indicating better separation.

4.2 Perfect CS does not imply accurate localization cues

As mentioned in the Introduction, studies have argued that the CS metric is not a robust indicator regarding localization performance. It can be shown mathematically that even perfect CS (i.e., R 12(ω)=R 21(ω)=0), does not in general imply small ITD/ILD/IACC errors. Let us assume that the target binaural signals satisfy the broadband ITD–ILD relation

Y L ( ω ) = g e j ω τ Y R ( ω ) , Mathematical equation: $$ \begin{aligned} Y_L(\omega )=g\,e^{-j\omega \tau }\,Y_R(\omega ), \end{aligned} $$(10)

and that the reproduced signals obey (8). Assuming perfect CS, W L (ω)=R 11(ω)Y L (ω) and W R (ω)=R 22(ω)Y R (ω), hence

W L ( ω ) W R ( ω ) = g e j ω τ R 11 ( ω ) R 22 ( ω ) · Mathematical equation: $$ \begin{aligned} \frac{W_L(\omega )}{W_R(\omega )} = g\,e^{-j\omega \tau }\,\frac{R_{11}(\omega )}{R_{22}(\omega )}\cdot \end{aligned} $$(11)

Therefore, CS constrains only the inter-channel leakage ratios (i.e., |R 12(ω)|/|R 11(ω)| and |R 21(ω)|/|R 22(ω)|), while the realized binaural cues depend critically on the diagonal ratio R 11(ω)/R 22(ω) (i.e., any imbalance in diagonal magnitude produces ILD error, and any imbalance in diagonal phase produces ITD error) (see Appendix B). Moreover, unless R 11(ω)/R 22(ω) is approximately a constant gain and a constant delay over the whole frequency analysis band, broadband interaural coherence will generally be reduced, degrading IACC.

Finally, for a geometrically symmetric loudspeaker arrangement, it can be shown that if R 12(ω)=R 21(ω)=0 at the displaced pose, then the diagonal ratio reduces to R 11(ω)/R 22(ω)=H 11(ω)/H 22(ω) ( Appendix C). Consider, for example, a leftward yaw rotation from a symmetric frontal pose. This movement lengthens the left-loudspeaker–left-ear path H 11(ω) and shortens the right-loudspeaker–right-ear path H 22(ω). Therefore, from the above relation and equation (11), the resulting imbalance between H 11(ω) and H 22(ω) directly changes the reproduced ITD and/or ILD, even if the off-diagonal leakage terms are zero.

Consequently, CS should be interpreted as a measure of crosstalk suppression rather than a direct proxy for localization performance under head movements.

5 Loudspeaker binaural reproduction system

For a loudspeaker-based 2 × 2 binaural reproduction setup as shown in Figure 1, the aim is to first estimate the transfer functions between each speaker – ear combination, forming the 2 × 2 matrix H(ω) of transfer functions, then to compute the inverse filter matrix C(ω) and finally to apply the binaural target signals. These processes are described below.

5.1 Path estimation and equalization

To estimate the impulse responses (IRs) for each loudspeaker–ear path, we used the least mean squares (LMS) algorithm [34] with a white-noise excitation emitted by the corresponding loudspeaker. For each microphone–loudspeaker pair, the objective is to identify an L-tap FIR filter h(n) whose coefficients approximate the acoustic IR of that path. The LMS coefficient update is

h ( n + 1 ) = h ( n ) + μ x ( n ) [ d ( n ) h ( n ) x ( n ) ] , Mathematical equation: $$ \begin{aligned} \mathbf h (n{+}1) = \mathbf h (n) + \mu \,\mathbf x (n)\,\big [d(n) - \mathbf h ^{\top }(n)\,\mathbf x (n)\big ], \end{aligned} $$(12)

where h(n)∈ℝ L is the FIR filter estimate at time index n, μ is the step size (here μ = 1 × 10−5), and ( ⋅ ) denotes transpose. The scalar d(n) is the recorded microphone signal for the considered path, and x(n)∈ℝ L is the tapped-delay vector formed from the loudspeaker excitation x(n) as x(n)=[x(n) x(n − 1) … x(n − L + 1)]. The corresponding instantaneous error is

e ( n ) = d ( n ) h ( n ) x ( n ) . Mathematical equation: $$ \begin{aligned} e(n) = d(n) - \mathbf h ^{\top }(n)\,\mathbf x (n). \end{aligned} $$(13)

Iterating (12) sample by sample yields the final FIR estimate h for each loudspeaker–ear path. The IRs obtained with this procedure for two loudspeaker arrangements with different spans (180° and 30° – as described in Sect. 6.1) under anechoic conditions are shown in Figure 4. Note that this LMS identification procedure uses a white-noise excitation and estimates only the loudspeaker–ear paths H(ω).

Thumbnail: Figure 4. Refer to the following caption and surrounding text. Figure 4.

Impulse responses (first 350 samples) for the anechoic simulation scenario for the 2 × 2 CTC system. Left: 180° loudspeaker span; Right: 30° span. Top two rows: right loudspeaker → right/left ears; Bottom two rows: left loudspeaker → right/left ears.

After the identification of all the impulse responses the inverse of the transfer function matrix must be estimated. This can be achieved using a least squares procedure or frequency-based equalization. Consistent with the relative studies [1, 2] we have chosen to compute the inverse using the Tikhonov regularization procedure defined in equation (14).

C = [ H H H + λ I ] 1 H H e j ω m . Mathematical equation: $$ \begin{aligned} \mathbf C =\left[\mathbf H ^{H}\mathbf H +\lambda \,\mathbf I \right]^{-1}\mathbf H ^{H}\,e^{-j\omega m}. \end{aligned} $$(14)

Here, I is the identity matrix, and λ >  0 is the Tikhonov regularization parameter controlling the trade-off between inversion accuracy and speaker effort. The factor e jωm applies a modeling delay of m samples (corresponding to the half of filter length) to obtain a causal inverse filter. The inverse transfer–function matrix can then be used to filter the desired stimuli (i.e., a reference signal modified to infuse the target ITD–ILD information) to generate the loudspeaker excitations. Here, white noise was also used as the reference signal and multiple ITD – ILD combinations were generated from it.

The benefit of this equalization approach is that one can control the loudspeaker effort by controlling the parameter λ, which is meaningful for the various comparisons between different loudspeaker configurations. Specifically, less ill-conditioned loudspeaker geometries tend to require less loudspeaker effort to faithfully reproduce the target binaural signals [2, 36, 37], thus for a fair comparison, the parameter should be adjusted accordingly to ensure equal speaker drive effort for the topologies under study.

6 Numerical simulations and experimental description

6.1 Simulation setup

Three symmetric loudspeaker pairs were modeled, with the right loudspeaker at angle ϕ ∈ {0 ° , 45 ° , 75 ° } relative to the interaural axis and the left loudspeaker at 180 ° −ϕ, thus creating three span combinations 180°, 90° and 30°. Each loudspeaker was represented as a circular arc of diameter 0.16 m, located 2 m from the listener’s head. Three types of head displacements – frontal, lateral, and rotational – were examined (Fig. 5a) using the head model in Figure 5b. Figures 5c and 5d illustrate the 30° and 180° loudspeaker spans respectively.

Thumbnail: Figure 5. Refer to the following caption and surrounding text. Figure 5.

(a) Head movements and experimental measurement apparatus, (b) COMSOL head model and head displacement conditions. Blue denotes rotation from the center of the head, red denotes lateral displacement and black frontal displacement, (c and d) 2 × 2 30° and 180° span configurations. (c and d) correspond to the same room configuration RC1 and both speaker geometries are illustrated for clarity. (e and f) correspond to room configurations RC2 and RC3 and only the 180° span is depicted for brevity. The TV and couch are depicted for visualization only and were not included in the simulations.

Time-dependent transient 2D FEM simulations were performed in COMSOL (Pressure Acoustics, Transient) to assess how head movements affect the measured ITD, IACC, and ILD. Simulations were carried out in the horizontal plane at ear height. The acoustic pressure field was discretized using quadratic Lagrange elements (second-order basis functions). To resolve the highest frequency of interest, f max = 1.4 kHz, the maximum element size was constrained to h max ≤ λ min/8, where λ min = c/f max and c = 343 m/s. The head was modelled as an ellipse assigned a hard boundary condition, ∂p/∂n = 0. For the free-field condition, the computational domain was surrounded with a perfectly matched layer (PML) at the outer boundary to suppress reflections. For the room simulations, wall boundaries were modelled using impedance boundary conditions as described in Section 6.3. The loudspeakers were prescribed an interior normal acceleration, which produced two-sided radiation into the adjacent half-spaces, yielding a bidirectional, dipole-like in-plane radiation pattern. No loudspeaker enclosure/baffle was included.

Regarding the basic – free-field simulations (Sect. 7), three target ITDs (0,  430,  750 μs) and five target ILDs (0,   ± 5,   ± 10 dB) were combined. For the 0 μs ITD case, only 0,   + 5,   + 10 dB ILDs were used due to symmetry, yielding a total of 13 ITD–ILD combinations. An ITD of 0 μs corresponds to a centered image, while 750 μs produces strong lateralization [38]. Although physically realized HRTFs do not exhibit contradictory ITD–ILD signs for a single source, perceptual studies show trading between the cues (ILD can partially counteract ITD and vice versa) [12, 30]. Therefore, counter-signed combinations were included for completeness.

In Section 8 the effects of additional factors were also assessed. Specifically, we also examined the effects of (i) reduced loudspeaker distance (1 m), (ii) increased head width (0.16 m), (iii) two constant shifts (−0.02 m) along the x-axis and y-axis, and (iv) six different reflective conditions as described in Sections 6.3 and 6.4.2.

All head-movement simulations and the experiment focused in the frequency range 200–1400 Hz, aligning with the band where humans are most sensitive to ITD changes. A sampling rate of f s  = 14 000 Hz was used, giving an ITD time resolution of Δt = 1/f s  ≈ 71.4 μs. The sampling rate was limited by the available processing resources (i.e., mesh size/memory considerations on the simulation computer). To enable sufficient modelling even under reflective conditions, the total length of the IRs was set to 5000 taps. In addition to full-band equalization using the complete IRs, two reflective conditions were also equalized using only the first 10 ms of each IR (≈140 taps), i.e., direct-path/early-response equalization. Restricting the equalization to the direct sound can be practical, since environmental changes (e.g., occupancy or air conditions) mainly affect the reverberant response and can require frequent re-calibration.

6.2 Simulated head model

To investigate various loudspeaker configurations while allowing precise head movements, a simplified two-dimensional (2D) head model (Fig. 5b) was developed in COMSOL. An ellipse was used to represent the head. The chosen axes were a = 0.15 m (head width) and b = 0.19 m (head length), based on reference measurements [39]. Note that the selected width 0.15 m denotes the interaural diameter with the pinnae excluded. Point probes were placed at the ear positions, 0.5 cm outside the outer ellipse along the interaural axis, and were moving with the head during simulations of different head poses. Thus, the distance between the two probe microphones was 16 cm.

The head model’s resulting ITD and ILD measurements for a point source positioned at azimuths 0°, 45°, and 75° relative to the interaural axis were compared against published data from Brungart and Rabinowitz’s KEMAR head model [38]. In Figure 6, the angle convention of the reference study is followed to enable a direct comparison.

Thumbnail: Figure 6. Refer to the following caption and surrounding text. Figure 6.

Comparison of free-field simulation results between the proposed 2D head model and Brungart and Rabinowitz KEMAR head, highlighting ITD and ILD measurements. Following Brungart’s and Rabinowitz’s convention, 90° denotes a source positioned at the side of the head, while 15° denotes an almost frontal source positioning.

Figure 6 shows ILD differences for the same azimuths under 500 Hz and 3000 Hz excitation, following the frequency examples reported by Brungart and Rabinowitz (note that for this validation run, the mesh size was refined to be appropriate for the broader bandwidth). To compute these values, a free-field transient simulation was set up in COMSOL with a miniature sound source (diameter 0.01 m). A white-noise stimulus (duration 0.4 s, bandwidth 100–4000 Hz) was rendered for the three angles and three source distances. For each source position, the ITD between the two binaural probes was computed over the full noise signal. For ILD, the FFTs of the two recordings were formed and the ILD spectrum (right–left level difference) was averaged using a 50 Hz bin width around the two frequencies of interest.

Figure 6 also compares the measured ITD values with KEMAR measurements under broadband excitation, showing close agreement. From Figure 6, it is evident, for example, that the natural ITD–ILD combination for a laterally placed source at 1 m at 500 Hz is approximately 7 dB and 750 μs. The deviations at 500 and 3000 Hz are mainly attributed to the simplified 2D approximation, with the absence of pinnae providing an additional limitation, particularly at 3000 Hz (note that the 3000 Hz case is included only to enable direct comparison with the reference study, as the present work evaluates ITD/IACC robustness only in the 200–1400 Hz band).

To provide a reference for the head-locked–world-locked analysis in Section 7.2 and to further validate the results, Figure 7 shows the computed ITD–ILD combinations of the publicly available SADIE KEMAR head dataset [40] for a source at 1.5 m across several yaw rotations, showing matching natural ITD–ILD pairings for a source at about this distance. For example, Figure 6 indicates that a source positioned at 1.0 m and 45° azimuth has an ILD of around 5.5–7.5 dB at 500 Hz, while the measured ITD is approximately 420 μs. In Figure 7, a source placed at the same azimuth and positioned slightly further at 1.5 m corresponds to an ILD of 5.8 dB (475–525 Hz band) and to an ITD of 470 μs (broadband 200–1400 Hz band), indicating close matching.

Thumbnail: Figure 7. Refer to the following caption and surrounding text. Figure 7.

ILD–ITD extracted from KEMAR head measurements (SADIE database) for a rotating source located 1.5 m from the center of the head. ILDs were obtained from spectral level differences averaged over 200–1400 Hz (blue) and 475–525 Hz (orange), while ITDs were estimated from cross-correlation of band-limited impulse responses over the same frequency ranges. 90° denotes a source positioned at the side of the head.

6.3 Effect of reflections

Reflective properties were introduced by assigning an interior boundary condition (impedance) to the walls of the rectangular enclosure surrounding the loudspeakers and the head, as shown in Figures 5c5f. In total, three room geometries were simulated. Figures 5c and 5d indicates room configuration 1 (RC1), Figure 5e indicates RC2 and Figure 5f indicates RC3. For condition RC1, simulations were run for two wall impedances Z = 1000 Pa s/m (light reflections) and Z = 5000 Pa s/m (strong reflections). For condition RC2, only the impedance Z = 1000 Pa s/m was simulated. Regarding the condition RC3, only the impedance Z = 5000 Pa s/m was simulated.

Thus, in total four simulation conditions were chosen to represent different combinations of symmetric and asymmetric configurations, and two types of reflective wall conditions: RC1_1000, RC1_5000, RC2_1000, and RC3_5000.

Reverberation time was estimated for the RC1 configuration with the interrupted-noise method following ISO 3382 [41]. The two reflective boundary conditions (Z = 1000 Pa s/m and Z = 5000 Pa s/m) yielded RT60 ≈ 0.09–0.12 s and RT60 ≈ 0.32–0.44 s, respectively across the 1/3-octave bands in 200–1400 Hz. The strong-reflection case is consistent with the EBU recommended reverberation-time range for listening rooms [42].

Figure 8 displays the impulse responses obtained with the LMS procedure between a 75° loudspeaker and an ear microphone under various simulation conditions. Additionally, the impulse response of the analogous real room experiment as discussed in Section 6.5, is included for direct comparison. To show the various sound phenomena (i.e., wall absorption/reflection, wave interference, head scattering etc), Figure 9 shows a pressure map for a single time step in the simulation environment under light reflective conditions during a binaural reproduction scenario.

Thumbnail: Figure 8. Refer to the following caption and surrounding text. Figure 8.

Impulse responses between a loudspeaker positioned 75° from the interaural axis (i.e., the right loudspeaker of the 30° span configuration), and the right ear, obtained from (a) free-field simulation, (b) light reflective simulation (RC1), (c) strong reflective simulation (RC1), and (d) real-room experiment. Note: The experimental impulse response is time shifted to account for the soundcard’s additional buffer delay.

Thumbnail: Figure 9. Refer to the following caption and surrounding text. Figure 9.

Sound pressure 2D map at a time step, for the one listener 2 × 2 CTC system, for the 30° span configuration. The reproduction target is [ITD, ILD]=[750 us, 10 dB]. Light walls reflections (1000 Pa⋅s/m) condition. The simulations depict the various sound phenomena in action (e.g., wall reflections, interference patterns, and head scattering in the simulated room).

6.4 Simulations pipeline overview

6.4.1 Basic simulations

The main simulations of this study are presented in Section 7. In total, 13 × 3 × 3 × 4 + 13 × 3 = 507 simulations were run (13 ITD–ILD target pairs ×3 loudspeaker configurations ×3 movement types ×4 displacements, plus 13 × 3 for the initial pose). Simulations were performed utilizing the Python package mph allowing to (a) compute the signals in Python for the various loudspeakers excitations, (b) send the signals and the various parameters to COMSOL for simulation and (c) post-process the measured binaural signals to compute the various binaural metrics. An illustration of two simulations regarding the produced 2D sound pressure level (SPL) map around the head and the corresponding binaural signals for a target reproduction scenario [ITD = 0 μs, ILD = 0 dB] for the initial and a rotated head position can be seen in Figure 10.

Thumbnail: Figure 10. Refer to the following caption and surrounding text. Figure 10.

Effect of a 30° leftward head rotation on the recorded binaural signals and SPL maps for the ITD–ILD target [0 μs, 0 dB]. Free-field simulations. (a) Binaural recordings at the initial head position (IHP): The left- and right-ear signals overlap. (b) Room-scale SPL distribution for the IHP. (c) Recordings after a 30° leftward rotation: The left- and right-ear signals are separated, revealing interaural differences. (d) Local SPL map around the head at IHP. (e) Local SPL map around the head after rotation. Overall, the rotation increases both ITD and ILD relative to the IHP case.

Because the small span (30°) loudspeaker topology naturally requires higher drive effort due to reduced head shadowing (thus allowing more crosstalk residual energy between a speaker and the two microphones), the regularization variable λ was set so that, at the initial head position, the most demanding target in our evaluations (ITD = 750 μs and ILD = 10 dB) could be realized while maintaining a minimum interaural cross-correlation of ∼0.8 under this geometry. For the other loudspeaker geometries, λ was then chosen to yield the same drive effort, enabling a fair comparison.

Specifically, the power of the inverse filters for each loudspeaker i was computed for the 30° span configuration with [2]:

P i = 1 N k = 0 N 1 ( | C i 1 ( k ) | 2 + | C i 2 ( k ) | 2 ) , Mathematical equation: $$ \begin{aligned} P_i = \frac{1}{N}\sum _{k=0}^{N-1} \Bigl (|C_{i1}(k)|^{2} + |C_{i2}(k)|^{2}\Bigr ), \end{aligned} $$(15)

where C ij (k) denotes the k-th frequency-bin value of the corresponding row and column of the computed inverse filter matrix, and N is the number of frequency bins. Then for the other two loudspeaker geometries the λ parameter was set so that the power of the inverse filters matched the power of the 30° span geometry. The chosen λ values were [0.55 × 10−4, 0.65 × 10−4 and 0.8 × 10−4] for the corresponding spans [180°, 90° and 30°] respectively.

6.4.2 Complementary simulations

Section 8 describes the results of additional simulations examining other effects like the influence of speaker distance, head size, wall reflections, and the effect of constant x and y-axis head shifts (the shifts were introduced to examine how rotations will be influenced when the center of rotation is off center from the modeled condition). In total, 4 × 2 × 3 × 4 × 11 + 4 × 2 × 1 × 11 = 1144 simulations were run (4 ITD–ILD pairs × 2 spans × 3 movement types × 4 displacement steps × 11 conditions, plus 4 × 2 × 1 × 11 for the initial pose).

The last six conditions correspond to reflective simulations, the first four of which are as described in Section 6.3. In addition, the reflective simulations RC2_1000 and RC3_5000 were repeated with equalization of only the early part of the impulse responses (i.e., the direct-arrival/early-response portion before the later reverberant tail), resulting in total in six reflective conditions.

Again, to ensure a fair comparison between the two geometries, an equal effort regularization procedure (Eq. (15)) was followed. Specifically, a constant λ was used for all conditions under the 30° span configuration, while for each condition under the 180° span condition, the λ was modified to ensure equal effort for the two geometries under the same condition. The initial λ value was chosen so that both configurations had zero ITD error and an IACC error of less than 0.15 for the IHP under all conditions.

6.5 Real-room measurement methodology

The results were additionally validated by an objective experiment in a real room (Fig. 11).

Thumbnail: Figure 11. Refer to the following caption and surrounding text. Figure 11.

Head and loudspeaker placement relative to the room geometry for the 30° and 180° span cases. Red and green enclosing lines indicate the 30° and 180° span configurations respectively.

Specifically, two loudspeaker configurations were tested in a rectangular room (Fig. 11) with dimensions 5.0 m × 8.0 m × 3.3 m (width×length×height) and an average reverberation time RT30 = 0.8–1.0 s over the band of interest; specifically: 250 Hz : 1.08 s, 500 Hz : 0.96 s, 1 kHz : 0.84 s, 2 kHz : 0.86 s, measured at the head positions. The listener head center was positioned at (2.3,  4.0) m. Loudspeakers were placed symmetrically with respect to the head center; small asymmetries relative to the room walls thus remained. At the IHP, the interaural axis was aligned with the room width.

Two Shure WL183 microphones were used, powered by Micropower P5400 phantom supplies, connected to a MOTU UltraLite AVB sound card and a desktop computer. Each microphone was positioned at the concha of the corresponding ear. Custom loudspeakers with 0.16 m driver diameter and wooden enclosures (0.28 m × 0.24 m × 0.24 m) were used. All ADC/DAC processing was performed by the sound card. A Python sounddevice callback routine controlled audio I/O.

A subject (one of the authors) sat on a chair with a measurement apparatus (similar to Fig. 5a) placed in front to track positional changes. A fixed wooden rod attached to the subject’s head served as the motion reference. During measurements, minor head-position deviations (e.g., ±5° or ±0.005 m) did not measurably affect ITD or interaural cross-correlation, as the induced time shifts were small relative to the temporal resolution 1/f s  ≈ 71.4 μs.

7 Free-field simulations

Figures 1214 show the main free-field simulation results under three head-displacement types. Each figure reports the measured ITD, ILD, and IACC versus displacement for the three loudspeaker configurations. The results are discussed in relation to the geometric interpretation introduced in Section 3 and to the previously reported CS-based robustness conclusions. Thirteen ITD–ILD excitation pairs were examined, using target ITDs of 0 μs, 430 μs, and 750 μs as the basis set.

Thumbnail: Figure 12. Refer to the following caption and surrounding text. Figure 12.

Free-field simulation results showing changes in measured ITD, interaural cross-correlation, and ILD across three loudspeaker span configurations (180°, 90°, 30°) during frontal head displacements for 13 target ITD–ILD combinations. Target ITDs are 0, 430, and 750 μs. Loudspeaker–head distance is 2 m. IHP denotes initial head position. Condition A indicates perfect matching between the target ITD pattern and the measured ITD at the initial head position. Condition B indicates perfect target–measured cross-correlation matching. C indicates imperfect matching.

Thumbnail: Figure 13. Refer to the following caption and surrounding text. Figure 13.

Free-field simulation results showing changes in measured ITD, IACC, and ILD across three loudspeaker span configurations (180°, 90°, 30°) during rotational head displacements for 13 ITD–ILD combinations. Target ITDs are 0, 430, and 750 μs. Loudspeaker–head distance is 2 m. IHP denotes initial head position. Condition A indicates strong matching between the target ITD pattern and the measured ITD for the rotated head under the 180° span configuration. Condition B indicates poor target–measured ITD matching for the same rotated head under the 30° span configuration.

Thumbnail: Figure 14. Refer to the following caption and surrounding text. Figure 14.

Free-field simulation results showing changes in measured ITD, IACC, and ILD across three loudspeaker spans (180°, 90°, 30°) during lateral head displacements for 13 ITD–ILD combinations. Target ITDs are 0, 430, and 750 μs. Loudspeaker–head distance is 2 m. IHP denotes initial head position. Condition A indicates poor matching between the target ITD pattern and the measured ITD for the displaced head under the 180° span configuration. Condition B indicates good target–measured ITD matching for the same head position under the 30° span configuration.

Note that in Figures 1214, the initial no-displacement condition (initial head position, IHP; middle group of 13 bars) is identical across all three figures, as it corresponds to zero head displacement. This condition is highlighted in yellow color.

Figure 12 displays the simulation results for frontal head displacements. The ladder-shaped target ITD pattern (induced by the 13 ITD–ILD combinations) was achieved under all loudspeaker configurations for all frontal displacements, owing to perfectly symmetric changes in all the loudspeaker–ear paths during the movement. A degradation in IACC is observed for the 30° span even at the initial head position (0.0 m displacement (Fig. 12, condition C)). This reduction is attributed to reduced head shadowing, which increases crosstalk between channels when the loudspeakers are closely spaced frontally. It can be seen that the harder ITD–ILD target (750 μs, −10 dB) had an IACC of ∼0.8 as specified by the chosen regularization parameter. The cross-correlation degradation depends on the target. Naturally higher ITD–ILD targets require greater array effort, leading to larger quality and localization losses. The 30° span configuration also failed to realize the larger ILD differences effectively even at the IHP condition.

The localization results agree with previous CS results which indicate significant robustness of all loudspeaker topologies under frontal displacements. However, the results reveal that the small span 30° configuration can create more spatial blur. Thus, regarding frontal displacements for the best localization performance the large span 180° configuration is preferable. The IACC degradation under the small span topology is an early indication that the ITD and IACC metrics can provide more information regarding localization performance. For example, in [1] the CS and the relative and absolute sweet spot metrics did not identify a measurable difference across different span topologies for this displacement type.

Figure 13 presents the rotational–displacement simulation results. In the 180° span case, ITD remained relatively stable for rotations up to ±30° because the left–right loudspeaker–ear paths changed symmetrically (e.g., Fig. 13, condition A). For this geometry the rotation center aligns with the midpoint of the line connecting the two loudspeakers, thus both direct and crosstalk paths increase or decrease by the same amount and the ITD is insensitive (Sect. 3). For larger rotations (±60°), as both ears gradually receive more similar signals (e.g., the left loudspeaker’s waveform reaches both ears, and analogously for the right), high ITD–ILD targets cannot be maintained.

For all ILD = 0 dB targets, the required ITD (either 0 μs, 430 μs, or 750 μs) was mostly achieved in the 180° span configuration even under a ±60° rotation (with small degradation under the highest rotations), while the ITD = 0 μs, ILD = 0 dB combination yielded perfectly stable interaural cross-correlation for all rotations (in this case, the least-squares solution reduces to both loudspeakers emitting the same signal). In total, this configuration is beneficial for head-locked reproduction scenarios (i.e., a sound image following the head) under head rotations.

The 30° span configuration produced substantially less stable ITD (e.g., Fig. 13, condition B) and ILD. As described in Section 3, a 30° leftward rotation lengthens the left loudspeaker-ear path while shortening the right, yielding asymmetric times of arrival (ToA) at the ears and ITD deviations. Here, the ITD decreases or increases proportionally with the rotation direction, as does ILD which indicates that the topology is useful for a world-locked reproduction scenario.

The improved stability of the ITD cue under the 180° (large span) configuration is important as it could not be inferred from previous research findings. As also discussed in Section 3, previous CS based research [1] indicated that the loudspeaker span is not an important parameter regarding rotational displacements. This result was also the outcome of previous ITD based research [3], which paradoxically indicated the small-span configuration as better. Thus, this rather significant observation that was confirmed both theoretically and with simulations resulted in a more in-depth analysis in Section 7.3.

Finally, Figure 14 presents the simulation results for lateral head displacements, where the listener’s head translates along the interaural axis. IACC, ITD, and ILD deteriorated as lateral movement increased. The 30° span configuration showed the smallest ITD deviation during lateral motion, as expected from the (approximately) equal changes in the direct paths, yielding smaller left–right time-of-arrival (ToA) differences for the direct paths than in the 180° case.

By contrast, the 180° span exhibited pronounced ITD and cross-correlation deviations because lateral shifts introduce asymmetric path-length changes and hence notable ear-wise ToA differences. Additionally, the ITD decreased or increased disproportionally with the movement direction (e.g., for a natural world-locked scenario, since all the initial image presentations here are located frontally or to the right, a movement to the left should result in the right ear leading more (i.e., positive shift to the ITD) which is not the case here), indicating that the topology is not suitable for a world-locked reproduction scenario.

These results are in line with the general guidelines based on CS metrics [1] that indicated that smaller span is preferable for robustness to lateral displacements. The localization robustness of the small span topology to head movements was also confirmed by Takeuchi et al. [3].

In summary, the simulations align with the theoretical observations in Section 3 that for 2 × 2 CTC under free-field conditions with symmetric loudspeaker placement, the 30° small span geometry is more robust to lateral head displacements, whereas the 180° large span geometry better tolerates rotational and frontal head movements in terms of ITD and IACC stability. Whether ITD stability under movement is desirable however, depends on the application, as discussed in Section 7.2.

Thumbnail: Figure 21. Refer to the following caption and surrounding text. Figure 21.

Absolute error maps (white = zero error) under head displacements for multiple (11) conditions (columns): (a) baseline (speaker–head distance 2 m, rotation axis at head center, head width 0.15 m, free field), (b) speakers at 1 m, (c) −0.02 m y-offset, (d) −0.02 m x-offset, (e) larger head width (0.16 m), (f) RC1_1000, (g) RC2_1000, (h) RC1_5000, (i) RC3_5000, (j) RC2_1000 (early IR equalization), (k) RC3_5000 (early IR equalization). Subplots A and B depict ITD errors, C and D depict IACC errors and subplots E and F depict ILD errors. Four target ITD–ILD pairs were simulated. IHP denotes initial head position.

7.1 Effect of the regularization parameter

The effect of the regularization parameter was examined under two edge cases representing a centralized (ITD = 0 μs, ILD = 0 dB) and a lateral (image to the right ear) (ITD = 750 μs, ILD = 10 dB) virtual auditory image.

From Figures 15a and 16a it is evident that the level of regularization affected minimally the ITD sensitivity for both topologies when the target excitation was [ITD = 0 μs, ILD = 0 dB]. In both cases the maximum ITD deviation between the different λ was around 75 μs corresponding to the available ITD resolution. Note that small non-monotonic bumps observed in ITD and ILD across λ (i.e., around λ = 0.00001) reflect a common regularization trade-off in CTC. If λ is too small the inverse can develop spectral peaks, whereas if λ is too large the deconvolution becomes less accurate [43]. This trade-off can result in a non-monotonic dependence of ITD/ILD error on λ, as small changes can redistribute residual crosstalk differently at the two ears, leading to variations in ITD and ILD for some displacements. Finally, the IACC error showed a rather significant increase with decreasing λ under lateral displacements for the 180° span configuration (Fig. 15a), and under rotational displacements for the 30° span configuration (Fig. 16a).

Thumbnail: Figure 15. Refer to the following caption and surrounding text. Figure 15.

Effect of λ across displacement types for two excitation conditions for the 180° span configuration. The x-axis ordering and background colors match Figures 16 and 21. IHP (yellow): initial head position; R (light blue): rotational displacements, L (light red): lateral (left–right) displacements, F (light gray): frontal (forward/backward) displacements. Top: absolute ITD error, middle: IACC error, bottom: absolute ILD error from the corresponding target.

Thumbnail: Figure 16. Refer to the following caption and surrounding text. Figure 16.

Effect of λ across displacement types for two excitation conditions for the 30° span configuration. IHP (yellow): initial head position; R (light blue): rotational displacements, L (light red): lateral (left–right) displacements, F (light gray): frontal (forward/backward) displacements. Top: absolute ITD error, middle: IACC error, bottom: absolute ILD error from the corresponding target.

Considering the [ITD = 750 μs, ILD = 10 dB] target, the 180° span again showed very low sensitivity to the regularization level in terms of ITD (Fig. 15b), indicating that for this loudspeaker geometry, ITD robustness is weakly affected by regularization. Lower drive effort, however, in this case led to increased blurring in IACC and ILD. Nevertheless, for such a large ITD target value, ILD deviations of up to 5 dB from the nominal 10 dB value are not expected to substantially affect perceived source direction [12, 30].

Still under the same target excitation, significant ITD sensitivity deviations depending on the regularization level, were instead observed for the ill-conditioned 30° span configuration (Fig. 16b) under both frontal and lateral displacements, indicating insufficient regularization when λ was set to 0.001. This directly shows that higher array effort is required by the small span configuration, especially under strong ITD–ILD binaural targets due to the increased crosstalk. Nonetheless, this configuration is very helpful under rotational displacements in world-locked reproduction scenarios as the ITD error naturally fluctuates depending on the direction of the rotation (e.g., in Fig. 16b, for a ±60° rotation the ITD error fluctuates between 1250 μs (opposite side image rotation) and 250 μs (i.e., the image stays on the same side and slightly rotates)). Finally, moving from λ = 0.0001 to λ = 0.00001 resulted in different effects during head rotations. Depending whether the rotation was aligned with the set ITD–ILD target, the increased regularization had either a negligible effect on all binaural metrics (30° and 60° cases) or a more measurable effect (−30° and −60° cases). Thus in this specific case, for a world-locked reproduction scenario, the regularization parameter should be carefully considered.

7.2 Effective perceptibility of head movements

To exemplify the observations from the simulations, Figure 17 provides estimates of the perceived virtual-image direction as a function of loudspeaker geometry for two extreme virtual source representations. The SADIE KEMAR angle-to-ITD mapping (Fig. 7) was used for the estimation. Specifically, we consider only a centered image (ITD = 0 μs, ILD = 0 dB; Figures 17a17d and 17i17l) and a strongly lateralized image positioned to the right (ITD = 750 μs, ILD = 10 dB; Figures 17e17h and 17m17p).

Thumbnail: Figure 17. Refer to the following caption and surrounding text. Figure 17.

Illustration of the perceived virtual sound image direction for the IHP (a, e, i, m), 30° leftward rotation (b, f, j, n), 0.03 m leftward lateral (c, g, k, o) and 0.03 m frontal (d, h, l, p) displacement. Dotted line: the perceived direction of the virtual source. Solid line: head-centred reference radius pointing to the listener’s right (ipsilateral 90° direction), rotated/translated with the head to depict the imposed motion. First row: 30° span configuration – central image [ITD = 0 μs, ILD = 0 dB]. Second row: 30° span configuration – lateral image [ITD = 750 μs, ILD = 10 dB]. Third row: 180° span configuration – central image [ITD = 0 μs, ILD = 0 dB]. Fourth row: 180° span configuration – lateral image [ITD = 750 μs, ILD = 10 dB].

For a simplified illustration, it is assumed that ITD = 750 μs corresponds to a heavily lateralized image, while ITD = 350 μs corresponds to an apparent rotation of approximately 35° of the perceived image, consistent with our head simulations and the SADIE measurements (Fig. 7). Furthermore, it is assumed that any occurring ITD ≥ 750 μs with a minimum of 2 dB ILD also corresponds to a lateralized image [11, 12, 30, 4446].

Figure 17 illustrates the effect of the various head movements to the perceived image direction depending on the loudspeakers geometry. For example, assuming a centered-frontal image at the IHP (Figs. 17a and 17i), a 30° rotation under the 30° span (Fig. 17b) causes an ITD increase (the sound reaches earlier the right ear) thus the auditory image stays almost still (it is only slightly rotated to the left). However, for the same movement, in the 180° span (Fig. 17j), the resulting ITD remains constant thus the image “moves” with the listener. Thus, the small span configuration is preferable for a world-locked centered image, where the image must remain fixed in external space as the head rotates, whereas the large span configuration is preferable for a head-locked centered image, where it must rotate with the head.

For lateral displacements and under world-locked reproduction, the 30° span configuration is advantageous in the central-frontal imaging scenario use-case (see Fig. 17c vs. Fig. 17k). The 180° span causes an unnatural shift to the perceived direction of the auditory image (Fig. 17k). In this case, for a frontal image, a leftwards lateral head displacement should cause a slight shift of the virtual image to the right, which is not the case. However, at least the 30° span manages to keep the image central (Fig. 17c).

For the strongly lateralized target image, however (see Fig. 17g vs. Fig. 17o), the 180° span configuration performs slightly better, since it more effectively retains the image on the right side. Even though the 30° span configuration is generally more robust to lateral displacements, this specific ITD–ILD target represents an exception. In this case, the 180° span can more easily preserve a strongly lateralized image, because the right loudspeaker is located exactly to the intended image direction. This illustrates that the optimal topology can depend on the target ITD–ILD combination. Nevertheless, the overall trend remains that the 30° span is preferable for lateral displacements.

7.3 ITD- versus CS-based robustness criteria

As discussed previously, strong CS does not guarantee accurate localization under virtual acoustics. Figures 18 and 19 illustrate further limitations of this metric. As can be seen in Figure 18, over the frequency region of interest (200–1400 Hz), CS exhibits pronounced frequency dependence, and in both rotational and lateral displacement cases, the inferred “best” topology can change with frequency. In particular, under both rotational and lateral displacement conditions, the CS curves suggest a topology reversal around the mid-band (approximately 600–800 Hz), implying that no single loudspeaker span is uniformly preferable when judged solely by CS. Comparing the two spans in Figure 18, it can be seen that the 180° span exhibit better CSL and CSR until 600–800 Hz, whereas the 30° span has lower channel separation for higher frequencies.

Thumbnail: Figure 18. Refer to the following caption and surrounding text. Figure 18.

Frequency-dependent left and right channel separation (CS) (in dB) under a 30° yaw rotation and 0.03 m lateral translation, for two loudspeaker topologies (180° and 30° spans). Legend values indicate averages over 200–1400 Hz.

Thumbnail: Figure 19. Refer to the following caption and surrounding text. Figure 19.

Frequency-dependent absolute ITD error under a 30° yaw rotation and a 0.03 m lateral translation, for two loudspeaker topologies (180° and 30° spans) and two representative cue targets: [0 μs, 0 dB] (top row) and [430 μs, 5 dB] (bottom row). Lines denote the deviation of the measured ITD from the corresponding target at each frequency bin. Legend values indicate averages over 200–1400 Hz.

By contrast, the ITD-based robustness assessment yields consistent topology-dependent trends across the whole frequency region. Figure 19 shows that the large span configuration is systematically more robust to yaw rotations, whereas the small span configuration is systematically more robust to lateral translations. Importantly, unlike CS, the ITD metric does not predict repeated topology reversals within the band where ITD cues are perceptually dominant. This behavior is consistent with the geometric interpretation in Section 3 that under rotations, the large span 180° topology preserves near-symmetric time-of-arrival changes for the dominant direct paths, whereas the ϕ = 30° span topology induces strongly asymmetric path-length changes and thus large ITD deviations.

These results also help clarify why some earlier conclusions about rotational robustness can differ from the present findings. For example, Parodi and Rubak [1] indicated that loudspeaker span plays a very limited role for rotational misalignment when robustness is assessed primarily via CS. Our results indicate that this conclusion does not generally extend to ITD-based robustness. In particular, as the span increases, we observed markedly improved ITD stability relative to smaller-span arrangements.

Additionally, while small and moderate spans (e.g., up to about a 60° span) may exhibit comparable ITD sensitivity, substantially larger spans can yield a clear and practically important improvement in ITD robustness under yaw rotations. More specifically, Takeuchi et al. [3], reported similar ITD sensitivity for two relatively small-span arrangements (10° and 60° span respectively). Figure 20 shows an analogous behavior for the same two spans (similar ITD slopes), while also demonstrating that extending the span beyond this regime produces substantially higher ITD robustness. In other words, the major separation between topologies is not necessarily visible when restricting attention to small spans, but emerges when larger-span configurations are also included.

Thumbnail: Figure 20. Refer to the following caption and surrounding text. Figure 20.

Effect of loudspeaker span on ITD robustness during yaw rotations for a central target (ITD–ILD = [0 μs,  0 dB]). Measured ITD is shown versus rotational displacement for four spans, (10°, 60°, ϕ = 45° and 180° spans). Increasing the span beyond 60° yields a pronounced reduction in rotation-induced ITD change.

In this regard, it is interesting to note that [3] reported that the small span configuration had smaller IACC error compared to the larger one, indicating that the smaller span configuration would be better. Interestingly, we observed a non-monotonic behaviour regarding the IACC in the same target condition [0 μs, 0 dB]. As can be seen in Figure 13, the IACC of the small span configuration (i.e., 30° span) is better than the 90° larger span configuration. That is, up to a loudspeaker span, the ITD robustness is slightly improved but the IACC error worsens as also indicated by Takeuchi et al. [3]. However, when the span is increased even more to 180°, both ITD and the IACC are improved.

Overall, these findings reinforce two practical implications. First, CS (and sweet-spot measures derived from CS) is not a reliable proxy for localization-relevant robustness, particularly under rotational perturbations where temporal-cue errors can dominate. Second, loudspeaker span can be a determining design parameter for rotational robustness when evaluated using ITD and interaural coherence metrics. Accordingly, it is suggested that perceptually grounded measures (e.g., ITD and IACC) should be reported alongside CS when evaluating binaural CTC systems, to assess not only inter-channel leakage but also the performance under virtual acoustics scenarios where the binaural cues are more relevant.

8 Additional considerations

Figure 21 summarizes additional factors under consideration across four ITD–ILD excitation conditions. All physical alterations were included in the inverse-filter design (i.e., impulse responses and inverses were recomputed for each reproduction condition for the IHP, except for the x and y-axis shift conditions).

The various conditions under examination are described in the caption of Figure 21. Notably, closer source distance had minimal effect on both configurations (compare ACEb with ACEa and BDFb with BDFa in Fig. 21).

Introducing a backwards a y-axis shift of −0.02 m (i.e., the IHP was shifted from the modeled position) had no effect. Effectively with this shift the ear-speakers geometry remained essentially unchanged and thus the left and right ToAs shifted similarly (see ABc vs. ABa). However, an x-axis shift of −0.02 m had significant implications under all head movements (see Bd vs. Ba) for the 180° span configuration as effectively this shift caused left right ToA asymmetries under this loudspeaker geometry (note that in this condition, the IHP also indicates an error due to the head shift – here the paths were not remodeled to match the new initial position as the goal was to assess the effect of combinatory displacements). However, the 30° span topology was significantly more robust, due to the retention of the ToA symmetry during lateral shifts.

Increasing the head width (condition e) had no impact on ITD sensitivity (compare Ae with Aa and Be with Ba). In general, considering the various conditions b–e, the cross-correlation and ILD sensitivity remained mostly stable in comparison with the initial condition a, under all displacements except of the condition d (x-axis movement).

The last six conditions f–k represent reflective conditions. A significant observation is that the presence of walls (conditions RC1, RC2 and RC3) did not significantly affect the ITD robustness in the 180° span configuration during displacements (compare Bf-k with Ba), due to the dominance of the direct paths, while slight improvements were also observed in some cases (e.g., under ±60° rotations, the error from the target ITD in conditions Bfgh is sometimes lower than the corresponding free field condition Ba). The general stability and improvement in ITD robustness for this configuration can be attributed to the symmetry of the reflective paths between the two speakers and the two ears, owing to the centered head position and the symmetric placement of the speakers within the room. This is apparent especially under perfectly symmetric reflective conditions (Bg) (i.e., centered head and a loudspeakers pair under a 180° span positioned totally symmetrically and centrally into the room). In this condition, under frontal and rotational head displacements, the reflective paths arriving at both ears change equally, thus the resulting ToAs at the left and right ears shift by the same amount stabilizing the resulting ITD. Furthermore, under the same conditions, the additional reflective paths can facilitate in the improvement of ITD robustness compared to free-field, due to the existence of multiple reflective paths that change to a lesser extent than the direct paths.

While the 30° span configuration also remained stable against frontal displacements under RC1 and RC2 reflective conditions (see conditions Af-k for frontal displacements), due to symmetric changes in the reflection paths, it was significantly more sensitive during rotations in addition to its inherent free-field sensitivity (compare Af-k with Aa for rotational displacements). In these cases, significant asymmetric reflections are created, amplifying the ToA differences between the right and left ears. Even though the reflections at the initial head position are also symmetric in this configuration, considering a rotational displacement, the reflective paths arriving at both ears change unequally (i.e., the left ear receives totally different reflections than the right ear resulting in uneven left -right ear ToAs). Lateral displacements also resulted in asymmetric changes in the reflection paths, but the effect was less detrimental.

The fourth reflective condition (i) corresponding to the only left–right asymmetrical configuration (RC3), had a significant effect on the frontal displacements for the 30° span configuration (compare Ai with Aa) in terms of ITD robustness, whereas the 180° span configuration had no degradation for the same condition (compare Bi with Ba). This result matches the experimental findings that were performed in the slight asymmetrical left–right room which also indicated that the ITD is more sensitive under frontal displacements under the 30° span configuration when left–right asymmetries are present. The near-orthogonality of the wavefronts in the 180° span case, together with equal left–right direct-path lengths (both conditions are satisfied during frontal motion), could help explain the better performance in this case.

Finally, two additional reflective-condition simulations were conducted using a shortened equalization window. In CTC, restricting the inversion to the early part of the binaural room responses is often used to reduce sensitivity to environmental changes (e.g., air conditions and occupancy), thereby reducing the need for frequent recomputation of the inverse filters. Specifically, the RC2 and RC3 analyses were repeated by equalizing only the first 10 ms of the impulse responses (i.e., a window extending ≈3 ms beyond the first IR peak). Accordingly, columns j and k should be compared against g and i, respectively.

ITD performance is similar across the corresponding conditions. However, IACC degrades with the shortened window, with the largest reduction in the highly asymmetric, strongly reflective case (column k) for both configurations, consistent with increased residual reflected energy. Thus, direct-path equalization is comparable to full equalization in symmetric, lightly reflective conditions (condition j vs. g), whereas in strongly reflective asymmetric conditions it yields reduced interaural coherence and less robust binaural cues (compare condition k vs. i), suggesting that more frequent filter re-adaptation may be required.

Prior work has reported benefits from accounting for early room effects regarding localization performance. Song et al. [9] incorporated estimated early reflections into the CTC transfer functions and reported reduced azimuth localization error in subjective tests. Thus, it remains a design choice whether and how often filter re-adaptation should be applied.

In general, for both topologies, even if in some cases the ITD robustness improved, the reflections significantly reduced the interaural cross-correlation under all conditions (compare CDf-k with CDa). This effect can affect the localization perception. Studies [24, 25, 47] provide evidence that an interaural cross-correlation of at least 0.4 is often required for sufficient localization.

Summarizing, for head-locked reproduction scenarios the large span 180° configuration can be considered more robust for rotational and frontal displacements (also note that the maximum cross-correlation error under these displacements with a maximum absolute displacement value of 0.03 m or 30° was 0.5), while the small 30° configuration was generally more robust for lateral movements in both anechoic and the reflective conditions under study.

9 Experimental results

This section reports the real-room measurements used to verify whether the main topology-dependent trends observed in the simulations also appear under experimental conditions. Figure 22 shows the absolute deviations of ITD, IACC, and ILD from the initial head position for two loudspeaker angles and four excitation types.

Thumbnail: Figure 22. Refer to the following caption and surrounding text. Figure 22.

Mean absolute deviations from target ITD, interaural cross-correlation, and ILD relative to the initial head position in the experimental setup, across two loudspeaker angles and four excitation combinations (S1–S4). Displacements are averaged over positive and negative directions.

The experimental results align with the simulations: the 180° span configuration retains the ITD cue more effectively during rotational displacements than the 30° span configuration, owing to the symmetric changes in the left–right direct paths as described earlier.

The small span topology was notably more sensitive to larger frontal displacements, consistent with the asymmetric left–right wall simulation (compare Ai with Aa under frontal displacements in Fig. 21). Reflection asymmetries were present in the experimental room due to slight off-center placement of the head and loudspeakers and left–right impedance mismatches (the right side was windowed, the left side concrete), leading to performance degradation especially for the more demanding ITD–ILD targets (e.g., the S1 excitation yielded optimal ITD robustness, whereas S4 showed the largest ITD sensitivity under frontal displacements). Finally, in agreement with the simulations, the 30° span topology was inherently more resilient to lateral displacements.

Finally, note that the large IACC error is in accordance with the simulations considering the asymmetric left–right wall condition (RC3) (Fig. 21, column i). Additionally, small head movements, measurement noise and imperfect equalization may also contribute to this result.

10 General discussion

This study showed that loudspeaker span significantly affects the robustness of ITD and IACC under head motion – two cues that are critical for precise localization. Findings indicate that the 30° (small span) configuration provides greater ITD stability against lateral translations, whereas the 180° (large span) configuration is more robust under rotational displacements and slightly preferable against frontal displacements. These tendencies were consistent across free-field simulations, reflective-room simulations, and experiments in a rectangular room. One of the most important findings was the improved ITD and IACC robustness of the large span 180° configuration contrary to the large ITD – IACC errors produced under rotational displacements by the small span configuration. The results indicated that the ITD robustness increases for larger spans, while the IACC can show a non-monotonic behavior.

Regarding the additional considerations, results support that speaker–head distance (1 m vs. 2 m) and head width had little impact on ITD robustness. A y-axis shift had minimal effect under all conditions. However, the ITD error increased when the rotation pivot point was x-shifted off-center in the large span configuration.

Under the reflective conditions RC1 and RC2, the 180° span configuration, considering the ITD robustness, benefited from nearly symmetric reflection-induced ToA shifts at the two ears, which increased ITD robustness. By contrast, the 30° span configuration suffered from asymmetric left–right ToA changes, especially during rotation, degrading ITD stability. While reflections can sometimes aid ITD robustness, this benefit is counterbalanced by a marked reduction in interaural cross-correlation. Thus future work could assess when these trade-offs are perceptually advantageous.

A degradation of the ITD is not always a negative factor considering the virtual imaging spatialization scenario. For world-locked virtual images a natural shift in the measured ITD signifies a world-locked perception of the sound object which is relevant for cinematic applications. However, for virtual reality applications ITD robustness may be more beneficial if the sound originates from the listener (head-locked auditory perception).

Finally, it should be noted that the results of this study could equivalently be interpreted as characterizing the perceptual effects of residual listener-tracking errors in head-tracked binaural systems. Such errors may for example arise from tracking latency or limited update rate. Within this perspective, the reported ITD and IACC deviations could provide topology-dependent bounds on localization errors.

11 Limitations

The numerical analysis was based on two-dimensional, transient finite-element simulations in the horizontal plane using a simplified rigid head model without pinnae. This approach captures the dominant horizontal-plane ITD mechanisms but neglects pinna-related spectral cues. While such spectral effects can be important for elevation and front–back perception, they are comparatively weak in the 200–1400 Hz band considered [48]. Therefore, omitting the pinnae is not expected to alter the central ITD-robustness mechanisms or the qualitative span-dependent trends reported.

Moreover, the current study imposed broadband ITD–ILD targets rather than reproducing full head-related transfer functions (HRTFs). Since our results show that performance depends on the targeted ITD–ILD combination, the absolute error magnitudes may differ under full HRTF reproduction, although the main span-dependent trends are expected to remain. Ultimately, dedicated perceptual experiments could verify localization and image stability under head motion for the proposed loudspeaker spans.

12 Conclusions

In summary, the main findings of this study are summarized below:

  • Theoretical and empirical evidence supports that CS is not a reliable predictor of binaural cue fidelity. Consequently, robustness rankings based on CS [1, 2] may fail to predict actual localization performance.

  • Under rotational (yaw) head movements, larger-span configurations yield monotonically greater ITD stability than smaller spans and naturally the 180° span yields the lowest ITD error, while small span 30° configurations exhibit very weak ITD stability. These findings were not previously indicated by CS – based sweet spot criteria.

  • Under rotational (yaw) head movements, the IACC does not necessarily vary monotonically with loudspeaker span across all target conditions. Nevertheless, the 180° span consistently yields the lowest IACC error.

  • Under rotational (yaw) head movements, while large-span topologies are stable in ITD for modest ITD–ILD targets, they are less effective under large ITD–ILD targets during the larger head rotations.

  • Under lateral head displacements, the results confirm previous findings that the small-span topology is overall the most robust in terms of ITD. However, for strongly lateralized targets, the 180° span topology can perform better.

  • During frontal head displacements, IACC changes minimally with displacement, but smaller-span topologies exhibit lower baseline IACC at equal loudspeaker effort. This implies potential losses in image unity and homogeneous spatial impression, suggesting that larger-span configurations should also be preferred for frontal displacements, contrary to the literature-preferred small span configuration.

  • The results indicate that the main observations that localization robustness under lateral movements is increased under small span configurations, while localization robustness under rotations is significantly better under large span configurations, also hold under multiple reflective conditions in rectangular rooms.

  • In some cases symmetrical reflections can help stabilize ITD, whereas asymmetrical reflective path changes may destabilize cue robustness. In untreated low-frequency rooms, these reflections are significant below 1400 Hz, where the ITD cue is most critical.

  • In virtual reproduction, the rendering goal (head-locked vs. world-locked) determines which topology is preferable. The same ITD deviation can be either detrimental or even preferable, depending on the intended behavior of the auditory scene.

Acknowledgments

The authors wish to thank the anonymous reviewers for their insightful comments.

Funding

This research received no external funding.

Conflicts of interest

The authors declare no conflict of interest.

Data availability statement

The data and code that support the findings of this study are available on request from the author.

Author contribution statement

Alberto Erspamer: Writing – original draft, Methodology, Software, Data curation, Investigation, Experiments. Dimitrios Mylonas: Methodology, Software, Data curation, Investigation, Experiments. Christos Yiakopoulos: Methodology, Review and editing, Supervision, Validation. Ioannis Antoniadis: Supervision, Project administration.

Ethics approval

No formal ethical approval was required as the experiments were non-invasive and the only human subject was one of the authors.

References

  1. Y.L. Lacouture Parodi, P. Rubak: Objective evaluation of the sweet spot size in spatial sound reproduction using elevated loudspeakers. The Journal of the Acoustical Society of America 128 (2010) 1045–1055. [Google Scholar]
  2. M.R. Bai, C.-C. Lee: Objective and subjective analysis of effects of listening angle on crosstalk cancellation in spatial sound reproduction. The Journal of the Acoustical Society of America 120 (2006) 1976–1989. [Google Scholar]
  3. T. Takeuchi, P.A. Nelson, H. Hamada: Robustness to head misalignment of virtual sound imaging systems. The Journal of the Acoustical Society of America 109, 3 (2001) 958–971. [Google Scholar]
  4. J. Rose, P. Nelson, B. Rafaely, T. Takeuchi: Sweet spot size of virtual acoustic imaging systems at asymmetric listener locations. The Journal of the Acoustical Society of America 112, 5 (2002) 1992–2002. [Google Scholar]
  5. Y. Lacouture Parodi, P. Rubak: Sweet spot size in virtual sound reproduction: a temporal analysis, in: Y. Suzuki, D. Brungart, Y. Iwaya, K. Iida, D. Cabrera, H. Kato, Eds.: Principles and Applications of Spatial Hearing, World Scientific, Singapore, 2011, pp. 292–302. DOI: https://doi.org/10.1142/9789814299312_0023. [Google Scholar]
  6. X. Ma, C. Hohnerlein, J. Ahrens: Concept and perceptual validation of listener-position adaptive superdirective crosstalk cancellation using a linear loudspeaker array. AES: Journal of the Audio Engineering Society 67 (2019) 871–881. [Google Scholar]
  7. M.F. Simón Gálvez, D. Menzies, F.M. Fazi: Dynamic audio reproduction with linear loudspeaker arrays. AES: Journal of the Audio Engineering Society 67 (2019) 190–200. [Google Scholar]
  8. M.-S. Song, C. Zhang, D. Florencio, H.-G. Kang: Personal 3D audio system with loudspeakers, in: 2010 IEEE International Conference on Multimedia and Expo, Singapore, 19–23 July 2010, 2010. DOI: https://doi.org/10.1109/ICME.2010.5583184. [Google Scholar]
  9. M.-S. Song, C. Zhang, D. Florencio and H.-G. Kang: Enhancing loudspeaker-based 3D audio with room modeling, in: 2010 IEEE International Workshop on Multimedia Signal Processing, Saint-Malo, France, 4–6 Oct 2010, 2010. DOI: https://doi.org/10.1109/MMSP.2010.5661990. [Google Scholar]
  10. F.L. Wightman, D.J. Kistler: Headphone simulation of free-field listening. I. Stimulus synthesis. The Journal of the Acoustical Society of America 85 (1989) 858–867. [CrossRef] [PubMed] [Google Scholar]
  11. D.R. Begault: 3-D Sound For Virtual Reality And Multimedia, 2nd Edn., NASA/TM–2000–209606, 2000. [Google Scholar]
  12. R.H. Domnitz, H.S. Colburn: Lateral position and interaural discrimination. The Journal of the Acoustical Society of America 61 (1977) 1586–1598. [Google Scholar]
  13. H. Lee, F. Rumsey: Level and time panning of phantom images for musical sources. Journal of the Audio Engineering Society 61, 12 (2013) 978–988. http://www.aes.org/e-lib/browse.cfm?elib=17075. [Google Scholar]
  14. B. Rafaely, V. Tourbabin, E.A.P. Habets, Z. Ben-Hur, H. Lee, H. Gamper, L. Arbel, L. Birnie, T.D. Abhayapala, P. Samarasinghe: Spatial audio signal processing for binaural reproduction of recorded acoustic scenes – review and challenges, Acta Acustica 6 (2022) 47. [CrossRef] [EDP Sciences] [Google Scholar]
  15. K. von Kriegstein, T.D. Griffiths, S.K. Thompson, D. McAlpine: Responses to interaural time delay in human cortex. Journal of Neurophysiology 100 (2008) 2712–2718. [Google Scholar]
  16. E.A. Macpherson, J.C. Middlebrooks: Listener weighting of cues for lateral angle: the duplex theory of sound localization revisited. The Journal of the Acoustical Society of America 111 (2002) 2219–2236. [CrossRef] [PubMed] [Google Scholar]
  17. J. Badajoz, J.-H. Chang, F.T. Agerkvist: Reproduction of nearby sources by imposing true interaural differences on a sound field control approach. The Journal of the Acoustical Society of America 138 (2015) 2387–2398. [Google Scholar]
  18. P. Majdak, B. Masiero, J. Fels: Sound localization in individualized and non-individualized crosstalk cancellation systems. The Journal of the Acoustical Society of America 133, 4 (2013) 2055–2068. [Google Scholar]
  19. J. Zheng, J. Lu, X. Qiu: Linear optimal source distribution mapping for binaural sound reproduction, in: Proceedings of the 43rd International Congress on Noise Control Engineering (INTER-NOISE 2014): Improving the World Through Noise Control, Melbourne, Australia, 16–19 Nov 2014, pp. 1–8. https://acoustics.asn.au/conference_proceedings/INTERNOISE2014/papers/p125.pdf. [Google Scholar]
  20. W. Tan, G. Yu, D. Rao: Influence of first-order lateral reflections on the localization of virtual source reproduced by crosstalk cancellation system. Applied Acoustics 202 (2023) 109165. [Google Scholar]
  21. A. Brughera, L. Dunai, W.M. Hartmann: Human interaural time difference thresholds for sine tones: the high-frequency limit. The Journal of the Acoustical Society of America 133 (2013) 2839–2855. [Google Scholar]
  22. J. Klug, M. Dietz: Frequency dependence of sensitivity to interaural phase differences in pure tones. The Journal of the Acoustical Society of America 152 (2022) 3130–3141. [Google Scholar]
  23. J.C. Middlebrooks, D.M. Green: Sound localization by human listeners. Annual Review of Psychology 42 (1991) 135–159. [Google Scholar]
  24. J. Blauert, W. Lindemann: Spatial mapping of intracranial auditory events for various degrees of interaural coherence. The Journal of the Acoustical Society of America 79 (1986) 806–813. [Google Scholar]
  25. B. Rakerd, W.M. Hartmann: Localization of sound in rooms. V. Binaural coherence and human sensitivity to interaural time differences in noise. The Journal of the Acoustical Society of America 128 (2010) 3052–3063. [Google Scholar]
  26. M.A. Akeroyd, J. Chambers, D. Bullock, A.R. Palmer, A.Q. Summerfield, P.A. Nelson, S. Gatehouse: The binaural performance of a cross-talk cancellation system with matched or mismatched setup and playback acoustics. The Journal of the Acoustical Society of America 121, 2 (2007) 1056–1069. [Google Scholar]
  27. T. Takeuchi, P.A. Nelson: Optimal source distribution for binaural synthesis over loudspeakers. The Journal of the Acoustical Society of America 112 (2002) 2786–2797. [Google Scholar]
  28. G. McLachlan, P. Majdak, J. Reijniers, H. Peremans: Towards modelling active sound localisation based on Bayesian inference in a static environment. Acta Acustica 5 (2021) 45. [CrossRef] [EDP Sciences] [Google Scholar]
  29. A. Carlini, C. Bordeau, M. Ambard: Auditory localization: a comprehensive practical review. Frontiers in Psychology 15 (2024) 1408073. [Google Scholar]
  30. W.M. Hartmann, B. Rakerd, Z.D. Crawford, P.X. Zhang: Transaural experiments and a revised duplex theory for the localization of low-frequency tones. The Journal of the Acoustical Society of America 139 (2016) 968–985. [Google Scholar]
  31. T. Lübeck, Z. Ben-Hur, D. Lou Alon, J. Crukley: Binaural reproduction of microphone array recordings with 2D video in mixed reality. Journal of the Audio Engineering Society 73 (2025) 461–470. [Google Scholar]
  32. A.D. Brown, F.A. Rodriguez, C.D.F. Portnuff, M.J. Goupell, D.J. Tollin: Time-varying distortions of binaural information by bilateral hearing aids. Trends in Hearing 20 (2016). DOI: https://doi.org/10.1177/2331216516668303. [Google Scholar]
  33. W.A. Yost, D.C. Tanis, D.W. Nielsen, B. Bergert: Interaural time vs. interaural intensity in a lateralization paradigm. Perception & Psychophysics 18, 6 (1975) 433–440. [Google Scholar]
  34. P.A. Nelson, H. Hamada, S.J. Elliott: Adaptive inverse filters for stereophonic sound reproduction. IEEE Transactions on Signal Processing 40 (1992) 1621–1632. [Google Scholar]
  35. F. Pausch, G.K. Behler, J. Fels: SCaLAr – A surrounding spherical cap loudspeaker array for flexible generation and evaluation of virtual acoustic environments. Acta Acustica 4 (2020) 19. [Google Scholar]
  36. E.Y. Choueiri: Optimal Crosstalk Cancellation for Binaural Audio with Two Loudspeakers. Princeton University, 3D3A Lab., 2010. [Google Scholar]
  37. M. Miyoshi, Y. Kaneda: Inverse filtering of room acoustics. IEEE Transactions on Acoustics, Speech, and Signal Processing 36 (1988) 145–152. [Google Scholar]
  38. D.S. Brungart, W.M. Rabinowitz: Auditory localization of nearby sources. Head-related transfer functions. The Journal of the Acoustical Society of America 106 (1999) 1465–1479. [CrossRef] [PubMed] [Google Scholar]
  39. M.D. Burkhard, R.M. Sachs: Anthropometric manikin for acoustic research. Journal of the Acoustical Society of America 58, 1 (1975) 214–222. [Google Scholar]
  40. C. Armstrong, L. Thresh, D. Murphy, G. Kearney: A perceptual evaluation of individual and non-individual HRTFs: a case study of the SADIE II database. Applied Sciences 8, 11 (2018) 2029. [CrossRef] [Google Scholar]
  41. International Organization for Standardization (ISO): Acoustics – Measurement of room acoustic parameters – Part 2: Reverberation time in ordinary rooms and other spaces. ISO 3382-2:2008, Geneva, Switzerland, 2008. [Google Scholar]
  42. European Broadcasting Union (EBU): Listening conditions for the assessment of sound programme material: monophonic and two-channel stereophonic. EBU Tech. 3276 (2nd Ed., May 1998), European Broadcasting Union, Geneva, Switzerland, 1998. [Google Scholar]
  43. O. Kirkeby, P.A. Nelson, H. Hamada, F. Orduna-Bustamante: Fast deconvolution of multichannel systems using regularization. IEEE Transactions on Speech and Audio Processing 6, 2 (1998) 189–194. [CrossRef] [Google Scholar]
  44. J.E. Mossop, J.F. Culling: Lateralization of large interaural delays. The Journal of the Acoustical Society of America 104 (1998) 1574–1579. [CrossRef] [PubMed] [Google Scholar]
  45. J. Encke, D. Reimann, W. Hemmert, F. Völk: On the role of interaural level differences in low-frequency pure-tone lateralization. Acta Acustica United with Acustica 104 (2018) 753–757. [Google Scholar]
  46. R.M. Baumgärtel, H. Hu, B. Kollmeier, M. Dietz: Extent of lateralization at large interaural time differences in simulated electric hearing and bilateral cochlear implant users. The Journal of the Acoustical Society of America 141 (2017) 2338–2352. [Google Scholar]
  47. L. Kong, Z. Xie, L. Lu, T. Qu, X. Wu, J. Yan, L. Li: Similar impacts of the interaural delay and interaural correlation on binaural gap detection. PLoS One 10 (2015) e0126342. [Google Scholar]
  48. M. Otani, T. Hirahara, S. Ise: Numerical study on source-distance dependency of head-related transfer functions. The Journal of the Acoustical Society of America 125 (2009) 3253–3261. [Google Scholar]

Cite this article as: Erspamer A. Mylonas D. Yiakopoulos C. & Antoniadis I. 2026. Interaural time delay during head movements under loudspeaker-based binaural reproduction. Acta Acustica, 10, 58. https://doi.org/10.1051/aacus/2026056.

Appendix A

Considering the two signals arriving at the two ears, if one is delayed by D 1 samples and the other is advanced by D2 samples (e.g., due to a head movement), the resulting signals can be expressed as xR, new[n]=xR[n − D1] and xL, new[n]=xL[n + D2].

Using the cross-correlation definition R a , b [ k ] = n = a [ n ] b [ n + k ] Mathematical equation: $ R_{a,b}[k] = \sum_{n=-\infty}^{\infty} a[n]\;b[n+k] $, the resulting cross-correlation can be expressed as

R x R , new , x L , new [ k ] = n = x R [ n D 1 ] x L [ ( n + k ) + D 2 ] = m = n D 1 m = x R [ m ] x L [ m + k + D 1 + D 2 ] = R x R , x L ( k + D 1 + D 2 ) . Mathematical equation: $$ \begin{aligned} \begin{aligned}&R_{x_{R,\text{ new}},\,x_{L,\text{ new}}}[k] = \sum _{n=-\infty }^{\infty } x_R[n - D_1]\; x_L[(n+k) + D_2] \\&\qquad \mathop {=}\limits ^{m=n-D_1} \sum _{m=-\infty }^{\infty } x_R[m]\; x_L[m + k + D_1 + D_2] \\&\qquad = R_{x_R,\,x_L}\!\big (k + D_1 + D_2\big ). \end{aligned} \end{aligned} $$

Thus, the resulting cross-correlation has a lag at (k + D 1 + D 2). That is, the measured maximum is advanced by D 1 + D 2 in comparison with the initial head position.

Appendix B

Channel separation (CS) quantifies suppression of inter-channel leakage (Eq. (9)), while localization-relevant binaural cues depend primarily on the diagonal ratio R11(ω)/R22(ω) (Eq. (11)). Hence, even perfect CS does not guarantee accurate ITD, ILD, or IACC. Below, we consider the limiting case of perfect CS, i.e., R12(ω)=R21(ω)=0.

If the diagonal terms differ by unequal pure delays, R11(ω)=ejωδL and R22(ω)=ejωδR, then equation (11) yields an ITD offset τ′=τ + (δL − δR) without any CS degradation.

If instead the diagonal terms differ only in gain, R11(ω)=a and R22(ω)=b, the resulting ILD error is 20log10|a/b| dB, again with perfect CS.

Appendix C

For a geometrically symmetric loudspeaker arrangement, the plant matrix at the design pose takes the form , and its Tikhonov-regularized inverse inherits the same structure. Specifically, H0HH0 + λI is a matrix with equal diagonals and equal off-diagonals, whose inverse is

( H 0 H H 0 + λ I ) 1 = 1 p 2 q 2 [ p q q p ] , Mathematical equation: $$ \begin{aligned} \left(\mathbf H _0^H\mathbf H _0+\lambda \mathbf I \right)^{-1} = \frac{1}{p^2-q^2} \begin{bmatrix} p&\quad -q\\ -q&\quad p \end{bmatrix}, \end{aligned} $$

again with equal diagonals and equal off-diagonals. Right-multiplying by preserves this structure, yielding , i.e., C 11 = C 22 = A and C 12 = C 21 = B.

At a displaced pose, R = H C, thus

R 11 = H 11 A + H 12 B , R 12 = H 11 B + H 12 A , R 21 = H 21 A + H 22 B , R 22 = H 21 B + H 22 A . Mathematical equation: $$ \begin{aligned} R_{11}&=H_{11}A+H_{12}B,\qquad R_{12}=H_{11}B+H_{12}A, \\ R_{21}&=H_{21}A+H_{22}B,\qquad R_{22}=H_{21}B+H_{22}A. \end{aligned} $$

If perfect CS were achieved at the displaced pose, then R 12 = R 21 = 0, which gives B = −H 12 A/H 11 and H 21 = H 12 H 22/H 11 (note that B/A = −H 12/H 11 = −H 21/H 22). Substituting these relations into the diagonal terms gives

R 11 = A ( H 11 2 H 12 2 ) H 11 , R 22 = H 22 A ( H 11 2 H 12 2 ) H 11 2 · Mathematical equation: $$ \begin{aligned} R_{11}=\frac{A(H_{11}^{2}-H_{12}^{2})}{H_{11}}, \qquad R_{22}=\frac{H_{22}A(H_{11}^{2}-H_{12}^{2})}{H_{11}^{2}}\cdot \end{aligned} $$

Hence, for H 11 ≠ 0 and H 11 2 H 12 2 Mathematical equation: $ H_{11}^{2}\neq H_{12}^{2} $, it follows that R 11/R 22 = H 11/H 22. Thus, following from equation (11), in this idealized case the reproduced binaural-cue ratio is governed by the imbalance between the two ipsilateral loudspeaker–ear transfer functions. Consequently, unequal gain or phase in H 11/H 22 directly produces ILD or ITD error, respectively, even when the channel-separation metric indicates perfect crosstalk suppression.

All Figures

Thumbnail: Figure 1. Refer to the following caption and surrounding text. Figure 1.

Direct and crosstalk signals arriving in both ears under a binaural reproduction scenario.

In the text
Thumbnail: Figure 2. Refer to the following caption and surrounding text. Figure 2.

Direct and cross path changes under (a) rotational and (b) lateral displacements for the large span 180° 2 × 2 listener configuration. The black dots correspond to the initial ear positions before a head displacement, and the orange dots to the displaced ear positions. Thick lines (stronger energy contribution) correspond to the direct paths, and thin lines to the crosstalk paths.

In the text
Thumbnail: Figure 3. Refer to the following caption and surrounding text. Figure 3.

Direct and cross path changes under (a) rotational and (b) lateral displacements for a 2 × 2 small span listener configuration.

In the text
Thumbnail: Figure 4. Refer to the following caption and surrounding text. Figure 4.

Impulse responses (first 350 samples) for the anechoic simulation scenario for the 2 × 2 CTC system. Left: 180° loudspeaker span; Right: 30° span. Top two rows: right loudspeaker → right/left ears; Bottom two rows: left loudspeaker → right/left ears.

In the text
Thumbnail: Figure 5. Refer to the following caption and surrounding text. Figure 5.

(a) Head movements and experimental measurement apparatus, (b) COMSOL head model and head displacement conditions. Blue denotes rotation from the center of the head, red denotes lateral displacement and black frontal displacement, (c and d) 2 × 2 30° and 180° span configurations. (c and d) correspond to the same room configuration RC1 and both speaker geometries are illustrated for clarity. (e and f) correspond to room configurations RC2 and RC3 and only the 180° span is depicted for brevity. The TV and couch are depicted for visualization only and were not included in the simulations.

In the text
Thumbnail: Figure 6. Refer to the following caption and surrounding text. Figure 6.

Comparison of free-field simulation results between the proposed 2D head model and Brungart and Rabinowitz KEMAR head, highlighting ITD and ILD measurements. Following Brungart’s and Rabinowitz’s convention, 90° denotes a source positioned at the side of the head, while 15° denotes an almost frontal source positioning.

In the text
Thumbnail: Figure 7. Refer to the following caption and surrounding text. Figure 7.

ILD–ITD extracted from KEMAR head measurements (SADIE database) for a rotating source located 1.5 m from the center of the head. ILDs were obtained from spectral level differences averaged over 200–1400 Hz (blue) and 475–525 Hz (orange), while ITDs were estimated from cross-correlation of band-limited impulse responses over the same frequency ranges. 90° denotes a source positioned at the side of the head.

In the text
Thumbnail: Figure 8. Refer to the following caption and surrounding text. Figure 8.

Impulse responses between a loudspeaker positioned 75° from the interaural axis (i.e., the right loudspeaker of the 30° span configuration), and the right ear, obtained from (a) free-field simulation, (b) light reflective simulation (RC1), (c) strong reflective simulation (RC1), and (d) real-room experiment. Note: The experimental impulse response is time shifted to account for the soundcard’s additional buffer delay.

In the text
Thumbnail: Figure 9. Refer to the following caption and surrounding text. Figure 9.

Sound pressure 2D map at a time step, for the one listener 2 × 2 CTC system, for the 30° span configuration. The reproduction target is [ITD, ILD]=[750 us, 10 dB]. Light walls reflections (1000 Pa⋅s/m) condition. The simulations depict the various sound phenomena in action (e.g., wall reflections, interference patterns, and head scattering in the simulated room).

In the text
Thumbnail: Figure 10. Refer to the following caption and surrounding text. Figure 10.

Effect of a 30° leftward head rotation on the recorded binaural signals and SPL maps for the ITD–ILD target [0 μs, 0 dB]. Free-field simulations. (a) Binaural recordings at the initial head position (IHP): The left- and right-ear signals overlap. (b) Room-scale SPL distribution for the IHP. (c) Recordings after a 30° leftward rotation: The left- and right-ear signals are separated, revealing interaural differences. (d) Local SPL map around the head at IHP. (e) Local SPL map around the head after rotation. Overall, the rotation increases both ITD and ILD relative to the IHP case.

In the text
Thumbnail: Figure 11. Refer to the following caption and surrounding text. Figure 11.

Head and loudspeaker placement relative to the room geometry for the 30° and 180° span cases. Red and green enclosing lines indicate the 30° and 180° span configurations respectively.

In the text
Thumbnail: Figure 12. Refer to the following caption and surrounding text. Figure 12.

Free-field simulation results showing changes in measured ITD, interaural cross-correlation, and ILD across three loudspeaker span configurations (180°, 90°, 30°) during frontal head displacements for 13 target ITD–ILD combinations. Target ITDs are 0, 430, and 750 μs. Loudspeaker–head distance is 2 m. IHP denotes initial head position. Condition A indicates perfect matching between the target ITD pattern and the measured ITD at the initial head position. Condition B indicates perfect target–measured cross-correlation matching. C indicates imperfect matching.

In the text
Thumbnail: Figure 13. Refer to the following caption and surrounding text. Figure 13.

Free-field simulation results showing changes in measured ITD, IACC, and ILD across three loudspeaker span configurations (180°, 90°, 30°) during rotational head displacements for 13 ITD–ILD combinations. Target ITDs are 0, 430, and 750 μs. Loudspeaker–head distance is 2 m. IHP denotes initial head position. Condition A indicates strong matching between the target ITD pattern and the measured ITD for the rotated head under the 180° span configuration. Condition B indicates poor target–measured ITD matching for the same rotated head under the 30° span configuration.

In the text
Thumbnail: Figure 14. Refer to the following caption and surrounding text. Figure 14.

Free-field simulation results showing changes in measured ITD, IACC, and ILD across three loudspeaker spans (180°, 90°, 30°) during lateral head displacements for 13 ITD–ILD combinations. Target ITDs are 0, 430, and 750 μs. Loudspeaker–head distance is 2 m. IHP denotes initial head position. Condition A indicates poor matching between the target ITD pattern and the measured ITD for the displaced head under the 180° span configuration. Condition B indicates good target–measured ITD matching for the same head position under the 30° span configuration.

In the text
Thumbnail: Figure 21. Refer to the following caption and surrounding text. Figure 21.

Absolute error maps (white = zero error) under head displacements for multiple (11) conditions (columns): (a) baseline (speaker–head distance 2 m, rotation axis at head center, head width 0.15 m, free field), (b) speakers at 1 m, (c) −0.02 m y-offset, (d) −0.02 m x-offset, (e) larger head width (0.16 m), (f) RC1_1000, (g) RC2_1000, (h) RC1_5000, (i) RC3_5000, (j) RC2_1000 (early IR equalization), (k) RC3_5000 (early IR equalization). Subplots A and B depict ITD errors, C and D depict IACC errors and subplots E and F depict ILD errors. Four target ITD–ILD pairs were simulated. IHP denotes initial head position.

In the text
Thumbnail: Figure 15. Refer to the following caption and surrounding text. Figure 15.

Effect of λ across displacement types for two excitation conditions for the 180° span configuration. The x-axis ordering and background colors match Figures 16 and 21. IHP (yellow): initial head position; R (light blue): rotational displacements, L (light red): lateral (left–right) displacements, F (light gray): frontal (forward/backward) displacements. Top: absolute ITD error, middle: IACC error, bottom: absolute ILD error from the corresponding target.

In the text
Thumbnail: Figure 16. Refer to the following caption and surrounding text. Figure 16.

Effect of λ across displacement types for two excitation conditions for the 30° span configuration. IHP (yellow): initial head position; R (light blue): rotational displacements, L (light red): lateral (left–right) displacements, F (light gray): frontal (forward/backward) displacements. Top: absolute ITD error, middle: IACC error, bottom: absolute ILD error from the corresponding target.

In the text
Thumbnail: Figure 17. Refer to the following caption and surrounding text. Figure 17.

Illustration of the perceived virtual sound image direction for the IHP (a, e, i, m), 30° leftward rotation (b, f, j, n), 0.03 m leftward lateral (c, g, k, o) and 0.03 m frontal (d, h, l, p) displacement. Dotted line: the perceived direction of the virtual source. Solid line: head-centred reference radius pointing to the listener’s right (ipsilateral 90° direction), rotated/translated with the head to depict the imposed motion. First row: 30° span configuration – central image [ITD = 0 μs, ILD = 0 dB]. Second row: 30° span configuration – lateral image [ITD = 750 μs, ILD = 10 dB]. Third row: 180° span configuration – central image [ITD = 0 μs, ILD = 0 dB]. Fourth row: 180° span configuration – lateral image [ITD = 750 μs, ILD = 10 dB].

In the text
Thumbnail: Figure 18. Refer to the following caption and surrounding text. Figure 18.

Frequency-dependent left and right channel separation (CS) (in dB) under a 30° yaw rotation and 0.03 m lateral translation, for two loudspeaker topologies (180° and 30° spans). Legend values indicate averages over 200–1400 Hz.

In the text
Thumbnail: Figure 19. Refer to the following caption and surrounding text. Figure 19.

Frequency-dependent absolute ITD error under a 30° yaw rotation and a 0.03 m lateral translation, for two loudspeaker topologies (180° and 30° spans) and two representative cue targets: [0 μs, 0 dB] (top row) and [430 μs, 5 dB] (bottom row). Lines denote the deviation of the measured ITD from the corresponding target at each frequency bin. Legend values indicate averages over 200–1400 Hz.

In the text
Thumbnail: Figure 20. Refer to the following caption and surrounding text. Figure 20.

Effect of loudspeaker span on ITD robustness during yaw rotations for a central target (ITD–ILD = [0 μs,  0 dB]). Measured ITD is shown versus rotational displacement for four spans, (10°, 60°, ϕ = 45° and 180° spans). Increasing the span beyond 60° yields a pronounced reduction in rotation-induced ITD change.

In the text
Thumbnail: Figure 22. Refer to the following caption and surrounding text. Figure 22.

Mean absolute deviations from target ITD, interaural cross-correlation, and ILD relative to the initial head position in the experimental setup, across two loudspeaker angles and four excitation combinations (S1–S4). Displacements are averaged over positive and negative directions.

In the text

Current usage metrics show cumulative count of Article Views (full-text article views including HTML views, PDF and ePub downloads, according to the available data) and Abstracts Views on Vision4Press platform.

Data correspond to usage on the plateform after 2015. The current usage metrics is available 48-96 hours after online publication and is updated daily on week days.

Initial download of the metrics may take a while.