Multimodal Fusion with LLMs for Engagement Prediction in Natural Conversation
Authors:
Cheng Charles Ma,
Kevin Hyekang Joo,
Alexandria K. Vail,
Sunreeta Bhattacharya,
Álvaro Fernández García,
Kailana Baker-Matsuoka,
Sheryl Mathew,
Lori L. Holt,
Fernando De la Torre
Abstract:
Over the past decade, wearable computing devices (``smart glasses'') have undergone remarkable advancements in sensor technology, design, and processing power, ushering in a new era of opportunity for high-density human behavior data. Equipped with wearable cameras, these glasses offer a unique opportunity to analyze non-verbal behavior in natural settings as individuals interact. Our focus lies i…
▽ More
Over the past decade, wearable computing devices (``smart glasses'') have undergone remarkable advancements in sensor technology, design, and processing power, ushering in a new era of opportunity for high-density human behavior data. Equipped with wearable cameras, these glasses offer a unique opportunity to analyze non-verbal behavior in natural settings as individuals interact. Our focus lies in predicting engagement in dyadic interactions by scrutinizing verbal and non-verbal cues, aiming to detect signs of disinterest or confusion. Leveraging such analyses may revolutionize our understanding of human communication, foster more effective collaboration in professional environments, provide better mental health support through empathetic virtual interactions, and enhance accessibility for those with communication barriers.
In this work, we collect a dataset featuring 34 participants engaged in casual dyadic conversations, each providing self-reported engagement ratings at the end of each conversation. We introduce a novel fusion strategy using Large Language Models (LLMs) to integrate multiple behavior modalities into a ``multimodal transcript'' that can be processed by an LLM for behavioral reasoning tasks. Remarkably, this method achieves performance comparable to established fusion techniques even in its preliminary implementation, indicating strong potential for further research and optimization. This fusion method is one of the first to approach ``reasoning'' about real-world human behavior through a language model. Smart glasses provide us the ability to unobtrusively gather high-density multimodal data on human behavior, paving the way for new approaches to understanding and improving human communication with the potential for important societal benefits. The features and data collected during the studies will be made publicly available to promote further research.
△ Less
Submitted 13 September, 2024;
originally announced September 2024.
Dosimetric Evaluation of a New Rotating Gamma System for Stereotactic Radiosurgery
Authors:
Huan Liu,
Ahmed Eldib,
Lili Chen,
Bin Wang,
Shidong Li,
Curtis Miyamoto,
CM Charlie Ma
Abstract:
Purpose: A novel rotating gamma stereotactic radiosurgery (SRS) system (Galaxy RTi) with real-time image guidance technology has been developed for high-precision SRS and frameless fractionated stereotactic radiotherapy (SRT). This work investigated the dosimetric quality of Galaxy by comparing both the machine treatment parameters and plan dosimetry parameters with those of the widely used Leksel…
▽ More
Purpose: A novel rotating gamma stereotactic radiosurgery (SRS) system (Galaxy RTi) with real-time image guidance technology has been developed for high-precision SRS and frameless fractionated stereotactic radiotherapy (SRT). This work investigated the dosimetric quality of Galaxy by comparing both the machine treatment parameters and plan dosimetry parameters with those of the widely used Leksell Gamma Knife (LGK) systems for SRS. Methods: The Galaxy RTi system uses 30 cobalt-60 sources on a rotating gantry to deliver non-coplanar, non-overlapping arcs simultaneously while the LGK 4C uses 201 static cobalt-60 sources to deliver noncoplanar beams. Ten brain cancer patients were unarchived from our clinical database, which were previously treated on the LGK 4C. The lesion volume for these cases varied from 0.1 cm3 to 15.4 cm3. Galaxy plans were generated using the Prowess TPS (Prowess, Concord, CA) with the same dose constraints and optimization parameters. Treatment quality metrics such as target coverage (%volume receiving the prescription dose), conformity index (CI), cone size, shots number, beam-on time were compared together with DVH curves and dose distributions. Results: Superior treatment plans were generated for the Galaxy system that met our clinical acceptance criteria. For the 10 patients investigated, the mean CI and dose coverage for Galaxy was 1.77 and 99.24 compared to 1.94 and 99.19 for LGK, respectively. The beam-on time for Galaxy was 17.42 minutes compared to 21.34 minutes for LGK (both assuming dose rates at the initial installation). The dose fall-off is much faster for Galaxy, compared with LGK. Conclusion: The Galaxy RTi system can provide dose distributions with similar quality to that of LGK with less beam-on time and faster dose fall-off. The system is also capable of real-time image guidance at treatment position to ensure accurate dose delivery for SRS.
△ Less
Submitted 9 October, 2022;
originally announced October 2022.