Visual IVRs for Low-literate, non-literate
and Low-tech-savvy Users
Abstract
Here we discuss about the limitations of the traditional IVR system for low-literate, non-literate and non-tech-savvy users. IVR is the only technology that can be reached to these set of users. But IVR is a new technology for these users. They need appropriate training to use the IVR. We suggest an extended form of IVR system called Visual IVR. Visual IVR means visual support or graphical support, on the users’ mobile phone, to an existing audio based IVR. Visual IVRs aim is to minimize the training time and to lessen the cognitive load and errors of non-tech-savvy users.
Keywords
Visual IVR for low-literate and non-tech-savvy users, Multimodal IVR, Limitations of IVR, text and images interfaces.
Introduction
Currently, Interactive Voice Response (IVR) systems are widely used for customer support in businesses like banking, retail, telecom, product support etc. Many of the current applications have been designed for urban, tech-savvy, educated users. However, the same applications are currently used by all users. A large number of less educated, less tech-savvy users find it difficult to use such applications.
CHI 2012, May 5–10, 2012, Austin, TX, USA.
Prof. Anirudha Joshi
IDC, IIT Bombay, Mumbai [email protected]
Nagraj Emmadi
IDC, IIT Bombay, Mumbai [email protected]
Prasad Rashinkar
Tata Consultancy Services, Mumbai
PradnyaMalandkar
IDC, IIT Bombay, Mumbai [email protected]
Abhishek Shrivastav
IDC, IIT Bombay, Mumbai [email protected]
Nitendra Rajput
IBM Research, India Research Lab, India
Saurabh Srivastava
IBM Research, India Research Lab, India
On the other hand, IVRs are the only interactive technology that can be used on all types of mobile phones, and thus reach a large number of users. IVRs thus provide an opportunity to build a much wider range of services and products than the ones currently available.
For the new users of technology, different kind of IVR systems need to be designed, taking into account their capability of understanding and usage habits etc. This is because; most of this population is not familiar with IVR technology. They need intensive training to use this technology.
Providing a visual interface to an IVR could help reduce the training time for users. And also make the users more involved in performing the tasks.
Prior Work
Early IVR systems had directed dialog interactions i.e. after each response the system would direct the user to a small or limited set of choices. User has to respond via pressing a key (touch tone navigation) or by verbally answering some statements like yes or no etc. (Automatic speech recognition).
Automatic Speech Recognition (ASR) technology is a speech recognizer, which processes sound which is being supplied as an input. ASR uses these sound patterns and compares against the examples of human speech to which it has been previously exposed. ASR interprets the input and makes a respective decision. In touch tone navigation user gives an input via pressing a key on his mobile phone producing a Dual Tone Multiple Frequency (DTMF) tone. The DTMF tones
are recognized easily by IVRs without any errors. Depending on the DTMF tone of the particular key pressed, the decision is made.
Implicit and explicit confirmation would be required to remove the possibility of error. The use of implicit confirmations reduces the number of turns taken by the user, which makes the interaction less cumbersome but it places the onus on the user to respond to an error in an expected manner. Explicit confirmations, on the other hand, compel the user to respond through direct instructions, but of course these result in a longer interaction. The process of recovery from errors should go beyond mere audio prompting and use
multimodality, i.e., combine audio with icons or images, to offer contextually-relevant help. This will help or guide the user to easily and gracefully recover from an error. [1]
Navigating IVR in order to reach the human agent is a common interaction experience. Often it is difficult and frustrating due to the sequential and fixed paced nature of IVR menus.This common interaction experience can be significantly improved by augmenting IVR menus with coordinated visual displays. The more complex the menu tree, the greater the advantage of visual
augmentation.Interruptions and multi-tasking (e.g. listening to an IVR menu and watching cricket) in real world situations can make navigating a IVR menu even harder since critical information in the IVR menu can come up when the user’s attention is focused elsewhere. [2]
IVRs: Advantages and Disadvantages
The biggest advantage of IVRs is that it can be accessed by a large population – here and now.Anybody with a mobile phone can readily use an IVR. In developing countries, mobile phone penetration is much higher than any other device or technology. For example, India has an estimated 100 million internet users [3] and 884.37 million mobile phone users [4]. Worldwide, there are about 2 billion internet users [3] compared to 6.5 billion mobile phone users.
Cohen et al talk about several other advantages of spoken language interfaces [5]. They say that spoken communication plays a big role in our everyday life as we spend a substantial portion of our waking hours engaged in conversations. Such communication is implicitly learnt at a very young age as compared to most other user interfaces (e.g. choosing an item from a tool bar), which need to be learnt explicitly. Largely, spoken communication is an unconscious activity, where conscious attention of the user is on the
meaning of the message, rather than the interface [5]. In spite of these advantages, however, IVRs have several usability problems. Firstly, voice is necessarily linear, which makes any IVR time consuming and uninteresting. In contrast, visual interfaces can be perceived in parallel and users can scan visuals much faster than they can hear voice prompts. Tufte says that a talk which communicates information at the speed of 100-160 words per minute is not an especially high-resolution method of data transmission compared to visual information [6].
Secondly, voice is ephemeral and so puts additional cognitive load as user needs to remember all options and information being spoken by the machine. As a result, only limited amount of information can be communicated in voice-based interfaces. For example,
in a banking IVR application, it may not be possible to communicate a long list of transactions. Reading out such a list may take time. Further, the user may not be able to remember all the information being read out. In contrast, visual elements are persistent – these can be accessed before the voice prompt begins and after it ends.
Thirdly, visual elements are much better than voice at subtly communicating abstract notions of the interface such as hierarchy and grouping of menus or percentage bars for communicating progress. While these elements can be communicated in audio, these will need to be much more explicit and intruding than visual elements. As a result, we found that people with less exposure to technology need explicit training before they can use these IVR interfaces [7].
Finally, in all current IVRs, applications using visual information are naturally ruled out. This not only includes applications that need visual information explicitly, such as organising and sharing photographs, but also applications that could enhance the usability by using visual elements, such selecting available seats from a floor plan.
Visual IVRs
By a Visual IVRs, we mean an interface in which the main function of voice prompts is carrying out a directed dialog with the users (as in traditional IVRs), while the main function of the visual is to highlight options that are being talked about in the voice prompt and to make the interaction non-linear and lasting. The visual choices are visible before they can be heard in the voice prompt and are retained after the audio has been played out (as in graphical interfaces).
In this paper, we differentiate visual IVRs from the usual graphical interfaces that also have audio (or multi-modal support). Audio support has been extensively used for visual interfaces to provide feedback. The sounds of clicks when the user selects a browser link, beep sounds at the start or end of a process, a warning sound of an error dialog box, or a computer programme reading out that the time is now 11 o’clock, are examples of such interfaces. Often, such audio feedback can be intrusive – but this could be an advantage as well as disadvantage. As a result, there has been a trend to use audios subtly and only as a secondary means of communication.
In a graphical interface supported by audio, the primary drivers of the user interactions are the visual elements. As against this, in visual IVRs, the primary drivers of user interactions are the voice prompts, particularly as long as the user is unfamiliar with the application. Gradually, as the user becomes an expert with the application, the visual elements may become the primary drivers.
If designed appropriately, Visual IVRs can potentially make the best of both modes of interaction. The voice prompts can still explicitly direct the users to use the interface. The visual elements will make the user interactions non-linear and non-ephemeral. As a result, even if the attention of the user was to waiver for a while, the user can reorient himself more easily with the help of visual elements. Visual elements can communicate large amount of information, including visual information such as photographs. The use of multiple media will bring about redundancy. The combination of voice and visuals could make the interfaces easier to learn, faster to use and less
error-prone. It can therefore minimise the need for training users.
Use of text may sound counter-intuitive particularly if the product is meant for low-literate users. However, literacy is not a black and white phenomenon – there are many shades of grey in between. Adult users often lose their literacy skills because of lack of practice. Using text in visual IVRs will provide such practice. Moreover, differentiating text-free user interfaces has been found to be not liked by the less-literate as it feels that they have been singled out. An example for text plus image and audio interface is Shree Ganesha. Shree Ganesha is a phone designed for illiterate and low-literate users. A number pad interface is used in the phone as numerical literacy is more widespread than textual literacy. This design extends the use of the number-pad alone as the primary interface for
navigating all the functions of the phone.
Fig: (a) Shree Ganesha Screen
The phone gives an audio instruction for every action to be taken and also gives an audio feedback after
completion of a particular action. The voice instructions and icons helped users to achieve their tasks. [8]
TAMA Visual
Treatment Advice by Mobile Alerts (TAMA) is an IVR system developed to provide healthcare information services to people living with HIV/AIDS (PLHA) who are also the potential non-tech savvy users and low literate [7]. TAMA helps PLHA to take their pills on time, by giving reminders on their mobile phones. TAMA also helps the PLHA by giving treatment advice for simple day to day symptoms like fever, vomiting etc. TAMA not only gives a reminder but also takes the feedback from the users. The menu structure used in TAMA for pill reminder is very simple and short one. The pill reminder menu was as follows.
1) If you have taken the pill press one.
2) If you have not taken the pill but are going to take it later then press two.
3) If you are not going to take the pill then press three. User has to press the keys accordingly. In such case if we provide a visual component on the phone which resembles the audio being played in IVR, user would understand the menu much faster and would respond with no hesitation and eventually require less training. The visuals for the above menu structure would be as follows:
Fig: (b) Fig: (c) Fig: (d) The figures b, c and d show visuals for TAMAs pill taking menu structure which would make the system more self-learning and will minimize the training time, errors, memory load, etc. of non-tech-savvy users.
Conclusions
For low-literate, non-illiterate and non-tech-savvy users, traditional IVR systems are difficult to learn or understand without any training. Even if the users are trained on traditional IVR they tend to make mistakes or errors, because of non-audible sound or interruption in surroundings or simultaneously working on other tasks etc.
Using the Visual IVRs, training time of the users can be reduced. Users would learn on their own by perceiving visuals on the screen. The task completion time with visual IVR would be less as compared to traditional IVR. Looping in the menus will be minimized as user has all the options upfront to select anyone, so user doesn’t have to listen to the same menu again and again. The visual IVR system solves the majority of the issues faced by the traditional IVR systems.
Developing a visual IVR would be a different technical challenge which has to be considered. The
synchronization of the audio on IVR and visuals on mobile phone is the major problem to be solved.
References
1. Aditi Sharma Grover, O.: Designing Interactive Voice Response (IVR) Interfaces:Localisation for low literacy users. (2009)
2. Min Yin, S.: The benefits of augmenting telephone voice menu navigation with visual browsing and search., 319-328 (2006)
3. Internet World Statistics: Internet Users in the World. (Accessed July 31, 2011) Available at:
http://www.internetworldstats.com/stats.htm
4. TRAI: Telecom Subscription Data as on 31st July, 2011. (Accessed September 9, 2011) Available at:
http://www.trai.gov.in/WriteReadData/trai/upload/Press Releases/837/Press_Release_July-11.pdf
5. Michael H. Cohen, J.: Voice User Interface Design, Page 8. Addison Wesley (2004)
6. Tufte, E.: The Cognitive Style of PowerPoint: Pitching Out Corrupts Within, Page 15. Graphics Press LLC (2006)
7. Prasad Girish Rashinkar, A.: Healthcare IVRS for Non-Tech-Savvy Users., 20 (2011)
8. Anirudha Joshi, N.: Shree Ganesha: The First Phone for Illiterate Users. (2008)