Recognition rather than recall in voice user interfaces
I believe it was Norman who said, “put knowledge in the world” instead of forcing the user to keep it “all in their head”. Thus lowering the cognitive load of the user and thereby increasing the usability and learnability of a given product.
Now take recognition rather than recall. Human memory is optimised for recognition while being pretty terrible at recall, so it’s much better to provide users with a familiar set of options that they recognise rather than having to remember it from scratch.
This much makes sense. But when you move away from interfaces on screens that can afford the space to place a lot of “knowledge” on the screen, and move towards screenless interfaces such as voice user interfaces, these principles require new interaction patterns to be kept the principles intact.
While screen-based interfaces can be verbose or filled with a lot of data that is "always right there”, voice-based interfaces don’t have any “static” display of information. So that creates a design constraint, where to “put knowledge in the world”? Products like Amazon Echo/Google Home/Siri can even be said to be precise or even “brevity-focused” in the way that they are trying to answer certain types of questions. I believe this may require the user to keep more knowledge in their head in some cases and also makes it more challenging for users to use the technology for a variety of queries. But much of human to human “conversation” isn’t sequential, it’s more like ping pong where one is reacting to what the other is saying. We see some of that in AI using deep learning that has “memory” (i.e. back propagation) and may be able to be more adaptive and may even have spotaneity.
With voice user interfaces, users may be looking for “a few good things”, but in my experience, users are often looking for “the one right thing”(i.e. purchase a Fiji water), but this may be due to the way we perceive a machine’s “intelligence” in voice technology and our history (paying your bill over the phone using a voice-based automated service ) with voice-based interfaces.
Notice how Amazon’s new product, “Amazon Look”, uses a combination of voice-based interaction but it eventually switches to screen based interaction to finish the job. This isn’t inherently bad since media always needs to be “displayed” but could Look offer voice feedback on what looks good? Instead of going to the app, could Look play the “sidekick” role and tell her/him that she looks better in the first outfit? The ultimate question is would the user value the voice more than the screen based version? Although voice and screen may both be able to use the same technology that predicts which outfit is better, what is the qualitative difference of how humans value a voice telling us to choose the 2nd option as opposed to a screen?
Much to do in these areas and I’ve decided to start focusing my research and design efforts in these areas.
My hope is to push these types of interactions forward and do innovative research in this area. I’m still studying voice user interfaces and hope to study more embodied user interfaces (human-robot interaction, VAMR) in the near future.