← Back to Research Papers

Functionality Review of Mobile Apps for the Tracking and Self-Management of Fatigue: Systematic Search in App Stores and Content Analysis.

Authors: Abdulla AD, Sas C, Doherty G
Journal: JMIR mHealth and uHealth
mental health psychology open access

Abstract

Measuring the human in human-robot interaction (HRI) contexts is critically important if we want the robots we design to meet the high expectations we have for them: to be useful, helpful, reliable, trustworthy, transparent, and so on. The human perspective within HRI can be measured with objective and subjective measures. Objective measures, such as task completion time or task performance, can provide valuable evidence regarding the quality of a robot system, interface, or interaction (; ; ; ). However, when a human interacts with a robot, it is also important to account for and measure their internal psychological experiences, such as their perceptions, emotions, and other latent processes (; ; ; ). These experiences are not directly observable and therefore cannot be captured through behavioral observation or objective metrics alone (). Subjective measures, such as scales, play a critical role in accessing these latent, unobservable constructs. Constructs are defined as underlying psychological characteristics or processes that are not directly measurable (). Subjective measures provide valuable insight into a person’s attitudes, thoughts, feelings, expectations, and perceptions, offering a window into the individual’s internal state that would otherwise remain hidden. As these constructs are inherently unobservable, subjective measures such as scales are necessary for gaining a complete understanding of the human perspective during HRI. In HRI research, scales are frequently used to capture the human perspective and play a crucial role in contributing to advancements in robotics and related technologies. Evidence for their widespread use comes from , which detailed a comprehensive analysis of self-report measures in proceedings from the Human-Robot Interaction (HRI) and Robot and Human Interactive Communication (RO-MAN) conferences (two of the primary publication venues for HRI research) over a 7 year period (2015–2021). A major takeaway from this investigation was that the majority of publications from these conferences (889/1,464, or 61%) incorporated scales as dependent measures in study designs. However, of particular importance to this paper, it was also revealed that the vast majority of those that included scales (815/889, or 92%) relied on custom scales to capture the constructs of interest. The term “custom” refers to the customization of either an existing scale or the creation of a scale that has been customized for a specific purpose (e.g., a researcher comes up with a handful of items they believe will capture their construct of interest); in either case, they are called “custom” scales because they have not been validated. Therefore, a straightforward conclusion from these results is that while scales clearly play a central role in HRI research, the widespread usage of non-validated scales raises concerns about the reliability and replicability of existing findings, as such measures are often not comparable across studies and have not undergone rigorous scale development and testing. More recently, investigated HRI conference proceedings from 2016 to 2020 (excluding alt.HRI and Late-Breaking Reports) that specifically reported using Likert scales. The authors reported that a majority of these scales (72.5%) were not appropriately developed or validated according to acceptable psychometric standards. Importantly, an additional 9.8% of scales used were custom scales. One potential reason researchers use custom scales is that they may be unaware of existing scales that meet their research goals. Since HRI is inherently multidisciplinary, relevant scales can originate from many different fields, such as robotics, psychology, sociology, human factors, human-computer interaction, ergonomics, or elsewhere. Taken together, this evidence for limited use of validated scales is concerning and underscores a critical issue in the field: the scale selection process in HRI research can and should be improved.