While it is well established that Natural Language Interaction (NLI) plays a vi- tal role in modulating and enhancing the quality of social interactions between robots and people (Belpaeme et al., 2012), the current state-of-the-art of Natu- ral Language Processing (NLP) still suffers from notable short-comings that can significantly impact HRI in adverse manners (Mubin et al., 2009). For example, while Automatic Speech Recognition (ASR) has gradually become a reasonably robust and common place technology2, facets of NLP such as Natural Language
Understanding, Dialogue Management and Natural Language Generation still re- main challenging tasks. Furthermore, given the serial “pipeline” nature of NLP, there is very little room for error, and when errors do occur they quickly propagate and often lead to breakdowns in NLI, such as incorrect, or worse, no responses from the system, both of which are uncomfortable for users (Shiwa et al., 2009; Lee et al., 2010). This makes facilitating natural language in current robots a challenging and cumbersome task.
Strategies stemming from NLP for coping in situations where NLP might fail include constraining and scripting interactions and dialogues, narrowing the scope of user responses (e.g. Lohse et al. (2008b)), or employing a set of general purpose responses to try and catch the failing interaction (e.g. Lison and Kruiff (2009)). These strategies do have their limitations and inherent risks however, as incorrect
2Robust and common place, in this case, refers to the fact that ASR technology is now a common feature on most smart phones and tablets, and in some cases, games consoles. However, noisy environments, multiple speakers and speaker variation such as accents and age (to name a few), still pose problems.
or repetitive linguistic responses are quickly identified by users, often revealing the limitations of the system (Ros Espinoza et al., 2011). Such revelations tend to hamper the development of long-term, open-ended HRI, which is a long-term goal of the field (Belpaeme et al., 2012). In such situations it may be appealing to be able to disguise the limitations of the system from the user in some way to mitigate the overall negative effect, and if the problems extend to such a degree that NLI is no longer possible, perhaps to revert to a replacement modality, such that interaction can continue, albeit in a limited capacity. NLUs can potentially provide a solution in both cases. However, as we shall see in chapter 9, this potential use is built upon some fundamental assumptions regarding the use of NLUs with natural language, and the validity of these assumptions needs to be tested and confirmed.
While the short-comings in comparison to natural language are obvious, NLUs do have qualities that hold promise for HRI however. For example, utterances are not bound to a particular spoken dialect, thus their use in multi-lingual and cul- tural settings may be advantageous. Secondly, given that NLUs hold/communicate little semantic content, there is generally a lower need to process semantic infor- mation from user speech, thus settings that pose challenges for sensory equipment and technologies such as microphones and NLP can be considered less problem- atic3. Furthermore, less parsing and processing of semantic content results in
fewer delays in agent response times, helping bring the interaction closer to real time and aiding the fluidity of the vocal exchanges which has been shown to be crucial for HRI (Shiwa et al., 2009). Finally, as NLUs are generally considered to hold less semantic content (with less need for a robotic system to consider seman- tic content), the burden of interpretation lies with the user, the intelligent other, with their inherent understanding of situational context and natural tendency to anthropomorphise inanimate objects such as robots (Duffy, 2003), and treat them as socially competent (Reeves and Nass, 1996). Given this, the presence of an intelligent other may also be exploited to allow utterances to be used in far
3Such settings tend to be in real world environments that are far from the protected and “safe” laboratory environments.
less restricted scenarios, widening the range of potential application areas, where the person may project meaning onto the abstract sounds based upon how the interaction is unfolding. This notion is explored in chapter 8.
NLUs also have another, subtle, but powerful potential affordance - the ability to allow robotic designers to subtly manage user expectations. It is a common observation in HRI that as the sophistication of a robotic system increases, so too does the user’s expectations of the system, and thus the greater the risk that they discover the system’s limitations and disengage from the interaction (Ros Es- pinoza et al., 2011). This however can be circumvented through expectation setting where both information about a robot’s capabilities (e.g. vision, tactile sensing, speech recognition, etc.) and observable behaviour (e.g. reactive behaviour to input stimulus, and expressive displays, etc.) can be used as a tool to set user ex- pectations (Paepcke and Takayama, 2010). In theory, by employing NLUs rather than Natural Language, the robot designer is able to help keep the “bar” of ex- pectation low by producing a robot that does not risk engaging in open-ended NLI but can remain responsive to external stimuli and make expressive displays and engage in open-ended HRI. Again, gibberish speech provides a good tangible example of this through Kismet (Breazeal, 2002), where people were observed to readily engage in stimulating multi-modal interactions with the robot without the need to rely on natural language interaction. It is also worth noting that in such design philosophies, the ability for naive subjects to suspend disbelief (Duffy and Zawieska, 2012) is also used as a powerful tool, and is a aspect that could make the use of NLUs particularly useful during Child-Robot Interaction as children are observed to readily suspend disbelief and are very willing to engage in social interactions with robots (Robins et al., 2004; Belpaeme et al., 2012, 2013).
There are already a number of robotic systems, both fictional and non-fictional, that demonstrate how a variety of the qualities of NLUs can be applied to social robots. Moreover, there are also examples that stem for HRI research showing that not only are NLUs and gibberish a useful means of facilitating expressive vocal displays, but they also hold potential as a useful tool that be help advance
and support research into other areas of HRI in general. These are outlined here.
2.2.1
Utility as a tool in broader HRI research
As is shown later in this chapter, both NLUs and gibberish speech have the potential to be used beyond a tool for creating and animating expressive robots, but also as a tool for studying affective expression through sound and speech more generally. However, this section serves to point out that NLUs and gibberish speech have properties that make them very appealing as tools to be used in other areas of HRI also. The Kismet robot (Breazeal, 2002) is prominent example of how gibberish speech can be used as a means of vocal expression in a robot, but at the same time is used as a tool to help facilitate research into other areas of HRI simultaneously, such as evaluating the influence that affective models of the robot’s internal states can have on the observable behaviour of the robot.
For example, Chao and Thomaz (2013) have used gibberish speech in a similar manner with their robot, Simon. In this work, the focus of the research was on evaluating their computational model turn taking during multi-modal HRI. Their evaluation required subjects to interact with Simon, in a natural manner, and so they told subjects to teach the robot about a variety of different objects4.
As the focus of the work was on turn taking, the robot was required to engage in the interaction and make both visual gestures and audible vocalisations. In order to avoid having to implement an NLP system, which if it failed could have had adverse consequences on interactions, they implemented a gibberish speech system in the robot, in the same way as (Breazeal, 2002) and for the same reasons - to elicit natural behaviour and turn taking from the human, without the need to cater for increased complexity and risks that come with NLP and the use of natural language.
In research focused upon the physical, anthropomorphic design of robots and how this impacts the perception people have of a robot, Walters et al. (2007) used NLUs to facilitate vocal animation of a robot that was deemed to be “machine-
4In reality, the robot did not do any learning. By asking subjects to teach the robot about objects, they were subconsciously encouraged to behave and interact in a natural manner.
like”, as opposed to having a more anthropomorphic design. Again, in this ex- ample, NLUs have been used as a tool to facilitate vocal animation of a robot, in order to be able to study aspects of HRI that fall far beyond affective expression via sound. Another example of this use of NLUs is research in with the robot Keepon (Kozima et al., 2009), which is a robotic tool designed to be used with young autistic children, many of which are pre-verbal. In this cases, not only does the robot’s morphology not lead itself to the use of gibberish speech, but the use of natural language with pre-verbal infants serves little purpose and runs the risk of over complicating interactions.
In these examples, the benefits of NLUs and gibberish speech shines through clearly. The use of natural language in robots is currently cumbersome due to the limitations that the technology has, and if the research does not strictly re- quire natural language, but does require some form of vocal expression, NLUs and gibberish can be seen as attractive options that can be implemented with con- siderable ease in comparison to NLP. Furthermore, the examples above are only a select few which have actually not used natural language for vocal expression when they have not needed to. There are many examples of research experiments that have adopted the Wizard of Oz (WoZ) experimental method (Kelley, 1984; Riek, 2012), where there is (unknown to the subject) a human controlling aspects of the robot. Moreover, in the majority, the reason why WoZ has been used is to facilitate vocal expression and natural language5, which highlights two points:
firstly that NLP technology is not in a mature enough state where it can be im- plemented into robotic systems for state-of-the-art research, and secondly, that there is a growing body of HRI research that is becoming contingent upon vocal communication and natural language in order to progress, and WoZ is used as a means to circumvent this contingency. The particular problem with the latter point is that the robot systems that are ultimately used in this research are not autonomous systems, but rather are mock-ups. This in itself can be a limiting factor in the general progress toward creating fully autonomous social robots.
5A recent review by Riek (2012) found that approximately 70% of WoZ studies used this technique to facilitate vocal and natural language.
It is in this light that the use of NLUs and gibberish speech in HRI research can be considered as highly fruitful as it provides a means of creating vocally expressive robots that are autonomous and thus can be programmed to operate in a consistent manner, making their use in experimentation particularly useful, and they remove any bias that natural language may have. Something that is useful when exploring other modalities in a robot.