mindmap root((Human-AI Interaction)) Smell sensing Low-cost VOC time series SmellNet diversity and confounds GCMS contrastive transference Smell generation Twelve-note wearable basis Multimodal semantic-to-actuator control Feedback and personalization Neural sensory interfaces Trigeminal smell stimulation Spatial and breathing-aware control Safety-bounded learned waveforms Taste and temperature Intraoral taste modulation Scene-aware virtual food experiences Trigeminal heat and cool illusions Touch and physical intelligence High-resolution tactile gloves Fast reactive robot feedback OpenTouch synchronized demonstrations Human-centered deployment Closed sensing and action loop User agency and transparency Safe constrained actuators

The course closes by connecting multimodal AI to a human feedback loop that includes senses beyond language, vision, and audio.

The final lecture shifts the course’s multimodal pipeline into direct interaction with people.
- Input: sense physical signals, including smell and touch.
- Model: align, fuse, and reason across heterogeneous representations.
- Output: generate sensory interventions such as aromas, tastes, or temperature illusions.
- Feedback: observe the user’s response and adapt over repeated steps.
This closed loop makes perception, generation, personalization, safety, and ethics parts of one system design problem.

Watch this section on YouTube


Low-cost gas sensors turn smell into a multichannel time series, but their accessibility comes with noisy and context-sensitive measurements.

Smell sensing can be represented as a synchronized time series:
- six coarse VOC channels;
- temperature, humidity, and pressure context;
- approximately 1 Hz sampling;
- a response curve produced as a substance approaches and leaves the sensor.
The cheap, portable hardware enables broad collection, but useful models must learn relative dynamics and resist environmental confounds.

Watch this section on YouTube


SmellNet broadens smell recognition through repeated, diverse measurements and reveals both class structure and collection confounds.

SmellNet treats dataset diversity as part of the modeling problem. Each example combines long sensor traces with image, text, and chemistry metadata, and repeated collection across environments prevents the model from equating a substance with one room or season. Low-dimensional projections are useful diagnostics: separation suggests recognizable signatures, while unwanted clusters reveal temperature, humidity, location, or breath confounds.

Watch this section on YouTube


Contrastive transference uses scarce high-quality chemistry measurements to improve abundant low-cost sensor representations.

This is co-learning by transference:
1. Encode portable sensor traces and precise GCMS profiles.
2. Pull paired representations together with a contrastive objective.
3. Use the geometry learned from the paired subset to organize the much larger unpaired sensor collection.
4. Deploy only the inexpensive sensor encoder.
Temporal differencing emphasizes changes rather than unstable absolute levels, and sequence-aware architectures better match the data’s dynamics.

Watch this section on YouTube


A general smell encoder can be adapted to targeted detection, although real-world composition makes apparently simple applications difficult.

Pretraining across many substances supports downstream specialization, such as a binary peanut versus no-peanut classifier. Yet preparation changes the observed gas mixture, and coarse VOC sensors detect emitted gases rather than exact ingredients. Robust deployment therefore needs compositional data, realistic dishes, and evaluation across cooking processes—not only clean laboratory samples.

Watch this section on YouTube


Scalable smell generation requires a compact basis whose mixtures can approximate many target aromas.

The smell generator mirrors a low-dimensional generative codec:
- choose a compact set of twelve base aromas;
- represent a target as twelve release durations;
- transmit that vector to a wearable dispenser;
- mix the timed emissions into a perceived composite.
Coverage, mixture quality, portability, and refill constraints determine the basis just as strongly as model accuracy.

Watch this section on YouTube


A multimodal language model can translate a food image and optional description into actuator commands for zero-shot aroma synthesis.

The model acts as a semantic-to-physical compiler:
(image + optional text + actuator vocabulary\rightarrow 12\text{-duration control vector}).
Language contributes ingredients or preparation details that pixels cannot show. Evaluation is perceptual rather than purely symbolic: a useful output may be one that forms a recognizable composite or evokes the intended experience, even when no unique ground-truth mixture exists.

Watch this section on YouTube


Iterative user feedback improves sensory similarity and allows the smell generator to learn personal preferences.

Sensory generation benefits from an interactive optimization loop:
1. Generate a base-mixture proposal.
2. Collect natural-language or scalar feedback.
3. Update the twelve actuator values.
4. Retain user-specific tendencies for future zero-shot predictions.
The alignment between smell and descriptive vocabulary is crucial: feedback is only useful when the model can map subjective words to controllable physical changes.

Watch this section on YouTube


Electrical stimulation of the trigeminal nerve offers a chemical-free route to induced smell, but precise control is physiologically and computationally demanding.

Chemical-free olfactory display replaces aroma canisters with carefully timed electrical pulses. Its control stack combines respiratory sensing, spatial and airflow modeling, bilateral stimulation, and a library of odor-associated waveforms. Replacing the lookup table with an AI generator could expand the odor vocabulary, but this is a high-stakes actuator: personalization, hard output bounds, clinical validation, and fail-safe shutdowns are prerequisites.

Watch this section on YouTube


Spatial smell generation borrows modeling ideas from graphics and other spatial media, yet it remains grounded in olfactory physiology.

Olfactory display is spatial: head pose, room airflow, and unequal left-right stimulation alter the perceived source. Techniques from computer graphics help model continuity and environmental transport, while controlled user studies remain the final test because perceived similarity cannot be inferred from the waveform alone.

Watch this section on YouTube


High-resolution taste modulation must act inside the mouth, where timing, space, saliva, and safety make hardware design unusually difficult.

An intraoral interface can change taste at the moment it is perceived instead of only altering the food beforehand. Timed, bilateral delivery enables phase-specific control—during chewing, between swallows, or in the aftertaste—but the mouth is a difficult deployment environment. Miniaturization, dosage limits, hygiene, latency, saliva variability, and individual calibration are core system requirements.

Watch this section on YouTube


Taste actuators can alter diet and virtual-reality experiences, while scene-aware models could replace today’s hard-coded interventions.

Taste modulation separates the physical prop from the perceived food. A small set of safe taste modifiers can transform one base item into several virtual experiences, reducing logistical complexity. Moving from scripted mappings to AI would allow context-sensitive control, but the model should operate over a constrained action library with explicit dose and timing limits rather than freely inventing chemical commands.

Watch this section on YouTube


Trigeminal temperature illusions create sensations of heat and cold without changing the room’s physical temperature.

Perceived temperature can be manipulated through the same sensory pathway used for spicy heat and minty coolness. Instead of thermally conditioning an entire room, a wearable emits carefully controlled micro-doses near the trigeminal nerve. The approach is compact and energy efficient, but it changes perception rather than ambient temperature and therefore needs clear safety boundaries and user awareness.

Watch this section on YouTube


Embodied sensory interfaces expose both rich interaction opportunities and ethical concerns about filtering or replacing reality.

Human-facing multimodal systems close the loop through the body, so an incorrect output can have physical and psychological consequences. Useful designs should preserve opt-in control, expose when and how perception is being modified, bound every actuator, and avoid optimizing engagement at the cost of authentic environmental awareness. The unanswered deployment question is not only what can be generated, but when should the system remain silent.

Watch this section on YouTube


High-resolution piezoresistive gloves capture touch faster than ordinary video, giving robots a rapid feedback channel for manipulation.

Tactile sensing contributes both spatial detail and low-latency feedback:
- dense fingertip arrays localize contact;
- broader palm coverage captures distributed pressure;
- 100+ Hz sampling observes fast slips and impacts;
- synchronized vision supplies object and scene context.
For manipulation, vision plans and recognizes while touch supports rapid reactive control.

Watch this section on YouTube


OpenTouch and the lecture’s closing framework point toward multimodal models grounded in synchronized physical experience.

OpenTouch aligns three complementary views of action: what the person sees, where and how the hand moves, and what pressure the hand feels. Such synchronized demonstrations can train robots to imitate contact-rich behavior. More broadly, multimodal AI becomes an interaction loop—sense → represent and align → reason → act → observe human and world feedback → adapt—with safety and evaluation spanning every stage.

Watch this section on YouTube