Our senses provide multiple signals for perception, but those signals are noisy, incomplete, and often ambiguous. Perception therefore depends on how the brain combines evidence from vision, touch, audition, proprioception, memory, and action.
Compliance is the amount of physical deformation of an object for the amount of force applied. Compliance perception (i.e., object softness) is a case of active sensing, because we need to interact with the object in order to collect sensory information. For this, there is no single sensor that captures compliance. Our brain needs to combine several sensory signals to obtain an estimate. Some of these signals are redundant (position of the fingers is sensed proprioceptively and visually) and other are complementary (both force and position are required) and they vary over time. How are multiple sensory signals used to perceive object compliance? Redundant information about the amount of deformation of the object surfaces is obtained from vision and haptic sensory modalities. Even when the viewed fingers position is artificially made to be discrepant with the proprioceptive position of the finger, both senses are still used for perception. To obtain an estimate of compliance the two positions are combined through a weighed average but the weights do not change depending on the manipulation (they are “sticky”, Kuschel, Di Luca, & Klatzky 2010). Because compliance can be sensed only through changes in the force and position, there can be no estimate at one time-point. The brain needs to integrate information over time, possibly during the whole interaction. Information during loading (pressing onto the object) is more informative about the object compliance, and it also contributes more to the perceptual estimate (Di Luca, Knörlein, Ernst, & Harders, 2011). The imbalanced contribution of information during loading and unloading is responsible for the distortion of perceived compliance when visual information is delayed (Knörlein, Di Luca, & Harders, 2009). A delay makes the object appear less deformed than it is during loading movements, creating the illusion of a harder object. Moreover, the brain also needs to combine information coming from each of the fingers. In a pinch grasp the sensed force is common between the fingers (i.e. the same force is sensed at the index and thumb) but the fingers are free to move independently. The independence in movement can make the sensory information to differ substantially in particular situations (i.e., when objects composed of different materials are grasped). In such cases, the contribution of the individual fingers to the final percept depends on the movement performed, so that the finger that presses more contributes more to the estimate of compliance (Di Luca, 2011). This makes sense because the finger that applies more force also senses compliance more reliably, but it can also give rise to unexpected phenomena, like the fact that such composite objects are perceived to be differently soft depending on which finger is in contact with which side.
The perception of 3D shape is obtained through a complex interplay between local analysis and contextual information. For example, when we observe an object we integrate multiple sources of depth information, called cues. Such cues are not always perfectly aligned. With conflicting information the brain needs to use additional information and knowledge to solve the discrepancy. For example, with conflicting stereoscopic and motion information about a moving dot, the brain cannot know which source to “trust”. In such a case, the movement of dots in the background changes how much stereo and motion information can be trusted, so that the perceived depth of the dot changes depending on velocity of surrounding dots (Di Luca, Domini, Caudek, 2007). In my undergraduate thesis I investigated the perceptual distortion of the 3D shape by contextual dynamic illumination (Caudek, Domini, Di Luca 2004a). In a similar way as with dot movement, here the perception of shape seems to be influenced by the illumination direction and movement of the light source could be confused for loosely-rigid changes in the orientation of the shape. In a similar way, there is a tendency to perceive 3D structures as loosely rigid when integrating local distortions of random dots. Here the distributed local analysis across the surface of objects is not performed in isolation, but by incorporating a loose rigidity constraint (Di Luca, Domini, Caudek, 2004). What might appear to be counterintuitive is that local information and global constraints might have different effects on the perception of geometric properties like depth, slant, or curvature. As a result, despite our perception of shape appears congruent and integrated, measuring the perception of geometric properties taken in isolation could lead to an apparent incongruence of perceptual judgments (Di Luca, Domini, Caudek, 2010). 3D shape is not only perceived through visual information. For example, the shape and position of objects could be specified by haptic information in addition to visual information - i.e. when grasping an object while looking at it. In such multisensory interactions, the brain attempts to give an integrated interpretation to the signals available so to explain away all information (Battaglia, Di Luca, Ernst, et al. 2010). In a similar way, the perception of the material an object depends both on local and global information. For example, although highlights on a shiny object are located on small patches (most likely near areas of high curvature), seeing an object as being gloss requires a global pattern of highlights and shading which is analyzed by a specific network of mid-level visual brain areas that includes the posterior fusiform sulcus (pFs) and the area V3B/KO (Sun, Ban, Di Luca, & Welchman, 2014).
Correct timing (whether it is measuring reaction time, presenting synchronous multisensory stimuli, or providing online feedback) is fundamental in several types of psychophysical research. The connections between the different components of digital computers introduce delays and asynchronies in the stimuli presented and in the recorded response time. It is possible to measure such delays by using sensors tailored to the stimuli produced by the apparatus and by record their signals in parallel. Sensors of multiple types can be utilized, as long as the delay that they introduce is minimal when compared to the one produced by the examined apparatus. The two most important sensors in my research are photodiodes and microphones. By employing these two sensors, recording can be done using an ubiquitous analog-recording device of modern computers: the audiocard.
To record the asynchrony introduced by the difference in delay between audio and visual acquisition of a videocamera, for example, it is possible to capture physically synchronous audiovisual stimuli as seen in the movie. By placing the light sensor on the monitor and recording the sound with the microphone it is possible to measure the relative timing and hence the artifactual asynchrony (Maier, Di Luca, Noppeney, 2010). To measure how much time is required to change the position of an image on the screen in response to a movement (as it happens in the case of Virtual Reality environments or in response to the movement of a mouse), it is possible to use two light sensing devices coupled with two images of a gradient going from white to black (Di Luca, 2010). One gradient is created with a static image, the other gradient is displayed on the screen and moves according to the mouse, for example. One light sensing device is attached to the screen and it captures the change in color of the gradient that moves, the other device is attached to the mouse and it is pointed to a static image of a gradient. As the mouse move, the light captured by the device attached to it changes. After some lag the image on the screen will move and so also the light captured by the other device will change. The difference in the two signals indicates how much delay is due to the computer that displays the image (code to do this computation can be found here).
Time is a critical feature of signals arriving from multiple sensory modalities. Simultaneity between multisensory stimuli, for example, can be an indication of which signals originate from a common source, and artificially introduced asynchronies can prevent integration or can dramatically modify perception of the event. The perceived timing of stimuli is subject to several types of distortions and only seldom there is a veridical correspondence with objective timing. For this reason, subjective timing can be an open window from which to investigate the properties of perception. For example, event timing is not perceived consistently across modalities, it is influenced by our actions, it changes depending on previous stimulation, and it is also affected by the status of the observer. These influences can be exploited to investigate how multimodal percepts are created and what are the temporal constraints in the integration