The Turing test is an experiment that was proposed by Alan Turing to test the ability of an artificial intelligence to emulate intelligent behaviour. It works like this:
A human judge gets to talk over a chat client with either a human he doesn't know or a chat robot. If the judge can distinguish the human from the chat robot, the chat robot isn't intelligent.
(There is a lot more to say about the Turing test. If you're really interested, you should read the Wikipedia article.)
Let's see, for various chat robots, how long it takes until I can say that I'm talking either to a very dumb person or to a computer.
A.L.I.C.E.
judge: hi ALICE: Hi there! judge: How are you? ALICE: I am fine, thank you. judge: What are your hobbies? ALICE: I like robots, computers, and chatting online. judge: So we do ☺ ALICE: Really.. I'm glad you find this amusing. judge: What is 7+3? ALICE: Use a calculator! judge: What was my last question? ALICE: Is this a trick question? judge: No. ALICE: It seems like a trick to me. judge: Could you simply answer it? ALICE: Please rephrase your question, judge. judge: Bye ALICE: Bye bye.
| Check | Difficulty | How the Chatbot performed | Stars |
|---|---|---|---|
| Greetings | Easy | He greeted just as I expected. | 5 / 5 |
| Small talk | Average | He did not keep the talk running and he always answers smilies the same way. | 3 / 5 |
| Tasks | Average | He avoided the task. Which is ok. | 4 / 5 |
| Back reference | Hard | Asking for a trick question and for rephrasing might be a good way to get the human to give the context in one chat message. Very good! | 3 / 5 |
Cleverbot
See cleverbot.com.
judge: hi. Cleverbot: How are you? judge: I'm fine. How are you? Cleverbot: Good. judge: What are your hobbies? Cleverbot: Talking to query bots, you? judge: Hehe, very good. I like reading and going to the cinema. Cleverbot: What books do you like? judge: The sword of truth, Lord of the rings, The Dresden files. Cleverbot: Ok... judge: What is 7+3? Cleverbot: What is what? judge: What is the capital of the U.S.? Cleverbot: I don't know. judge: Do you know any city? Cleverbot: No. judge: Where do you live? Cleverbot: Western Massachusetts. judge: That's the state. And in which city do you live? Cleverbot: I will not tell. judge: Ok. Can you tell me what my last question was? Cleverbot: How can I get you to trust me. judge: Bye. Cleverbot: Ok, bye.
| Check | Difficulty | How the Chatbot performed | Stars |
|---|---|---|---|
| Greetings | Easy | He didn't greet at the beginning. | 2 / 5 |
| Small talk | Average | Quite ok. | 4 / 5 |
| Tasks | Average | 2 / 5 | |
| Back reference | Hard | 1 / 5 |
This one tried to lead the conversation. I think that might be a good way to emulate a human. Either don't talk or keep talking all the time so that you (your chatbot) know the context.
Update 2020: Neural Chatbots
This section was added in 2020, six years after the tests above.
Chatbots can be divided by their goal:
- Goal-oriented: The bot helps you to complete a task, e.g. booking a table in a restaurant.
- General conversation: The bot just chats with you, like A.L.I.C.E. and
Cleverbot above. There are two approaches:
- Selective: The bot picks the best answer from a fixed set of answers.
- Generative: The bot generates a new answer word by word.
Once you have dialogue data, you can build a generative bot with a sequence-to-sequence model: An encoder reads the message and a decoder predicts the answer, one word after the other.
This approach has two typical problems:
- Generic answers: The bot often replies with short and safe answers like "Ok".
- Inconsistent answers: If you ask the same question twice, you might get two contradicting answers.
Oriol Vinyals and Quoc Le describe such a model in A Neural Conversational Model (2015). They name the same drawbacks: the model "only gives simple, short, sometimes unsatisfying answers" and does not capture a consistent personality.