This week’s funny is again brought to you with the courtesy of xkcd. It demonstrates, quite adequately, what can go wrong in the process of interpretation of data.
Category: Research Design
Weekly Funny: The Dunning Kruger Effect – Again
It really has evaluation implications!
How Many Days Does it Take for Respondents to Respond to Your Survey?
At my consultancy we use SurveyMonkey for all our online survey needs. It is simple to use, reliable, and they are very responsive.
Their research and found that
The majority of responses to surveys using an email collector were gathered in the first few days after email invitations were sent, and
•41% of responses were collected within 1 day
•66% of responses were collected within 3 days
•80% of responses were collected within 7 days
The graph below maps the response rate against time.
The findings suggest that, under most circumstances, it would be best to wait at least seven days before starting to analyze survey responses. Sending out a reminder email after a week would probably boost the response rate somewhat.
SurveyMonkey also did some interesting analysis to answer questions like:
How Much Time are Respondents Willing to Spend on Your Survey?
Does Adding One More Question Impact Survey Completion Rate?
Go check it out!
Survey answers when you ask people to state the obvious
The following comic from doghouse diaries, and the results of an actual colour survey at xkcd tells you a little about the validity of surveys…
The write-up about the “male / female” categories and the controls they tried to implement for color blindness at the xkcd blog is also something worth reading.
Dunning-Kruger Effect and Evaluation
Justin Kruger and David Dunning published a paper in the Journal of Personality and Social Psychology (1999, Vol 77, No.6, 1121 -1134) and the term “Dunning Kruger effect” was coined. This is the abstract:
People tend to hold overly favorable views of their abilities in many social and intellectual domains. The authors suggest that this overestimation occurs, in part, because people who are unskilled in these domains suffer a dual burden: Not only do these people reach erroneous conclusions and make unfortunate choices, but their incompetence robs them of the metacognitive ability to realize it. Across 4 studies, the authors found that participants scoring in the bottom quartile on tests of humor, grammar, and logic grossly overestimated their test performance and ability. Although their test scores put them in the 12th percentile, they estimated themselves to be in the 62nd. Several analyses linked this miscalibration to deficits in metacognitive skill, or the capacity to distinguish accuracy from error. Paradoxically, improving the skills of participants, and thus increasing their metacognitive competence, helped them recognize the limitations of their abilities.
Errol Morris described how the following sad story about a guy called McArthur Wheeler, inspired Dunning’s scientific inquiry:
Wheeler had walked into two Pittsburgh banks and attempted to rob them in broad daylight. What made the case peculiar is that he made no visible attempt at disguise. The surveillance tapes were key to his arrest. There he is with a gun, standing in front of a teller demanding money. Yet, when arrested, Wheeler was completely disbelieving. “But I wore the juice,” he said. Apparently, he was under the deeply misguided impression that rubbing one’s face with lemon juice rendered it invisible to video cameras. If Wheeler was too stupid to be a bank robber, perhaps he was also too stupid to know that he was too stupid to be a bank robber — that is, his stupidity protected him from an awareness of his own stupidity.
What does this have to do with evaluators? All I suggest is that you should think a little about the Dunning-Kruger effect next time you ask people to rate their own competence level in a survey. You would not want to design such a survey without knowing that it is not a very smart thing to do, right?
Alternatively you might want to read an earlier post I did about it here.
Specificity and Sensitivity in tests
You are required to identify kids in need of remediation using a scholastic ability test.
If your test is highly specific, a low score will be able to identify everyone that requires remediation. – A lack of specificity indicates that some kids who require remediation are not identified.
If your test is highly sensitive, then a high score will clearly exclude anyone that does not need remediation.
SPIN and SNOUT are commonly used mnemonics which helps to remind us ofthe disticntion: A highly SPecific test, when Positive, rules IN disease (SP-P-IN), and a highly ‘SeNsitive’ test, when Negative rules OUT disease (SN-N-OUT)
Survey Design
I’m working on a retrospective pre-post competency survey, and needed to be reminded of some basics of survey design again.
I find Neuman’s chapter about survey design a good foundation: http://www.amazon.com/gp/product/0205457932.
Jane Davidson Makes some compelling arguments that suggest that we do need to think twice when we decide to make use of Likert type scales in surveys. http://genuineevaluation.com/breaking-out-of-the-likert-scale-trap/
She sugggests that rather than use the “strongly agree / disagree” type anchors, one could use evaluative terms like “inadequate / good” that might make the data easier to interpret.
Ive also found a list of possible Likert Scale Anchors that are most useful:
http://www.hehd.clemson.edu/prtm/trmcenter/scale.pdf
Why Competency Self Assessments are essentially flawed beyond redemption
I’m working on an evaluation to determine if a training programme for senior government managers makes a difference – In the competence level of the managers, and in the service delivery they are able to produce within their work context. We are tracking a wide evidence base about all of the participants, but the client is insistent that a competency self-assessment be included. We agreed, on the condition that this one piece of evidence will be used together with all of the other evidence we will be collecting throughout the study. The value that the competency self-assessment will add, is something we have debated in the team. The following entertaining post by Errol Morris, however, pretty much sums it all up:
http://opinionator.blogs.nytimes.com/2010/06/20/the-anosognosics-dilemma-1/
David Dunning, a Cornell professor of social psychology… wondered whether it was possible to measure one’s self-assessed level of competence against something a little more objective — say, actual competence. Within weeks, he and his graduate student, Justin Kruger, had organized a program of research. Their paper, “Unskilled and Unaware of It: How Difficulties of Recognizing One’s Own Incompetence Lead to Inflated Self-assessments,” was published in 1999.
Dunning and Kruger argued in their paper, “When people are incompetent in the strategies they adopt to achieve success and satisfaction, they suffer a dual burden: Not only do they reach erroneous conclusions and make unfortunate choices, but their incompetence robs them of the ability to realize it. Instead, …they are left with the erroneous impression they are doing just fine.”
It became known as the Dunning-Kruger Effect — our incompetence masks our ability to recognize our incompetence.
There have been many psychological studies that tell us what we see and what we hear is shaped by our preferences, our wishes, our fears, our desires and so forth. We literally see the world the way we want to see it. But the Dunning-Kruger effect suggests that there is a problem beyond that. Even if you are just the most honest, impartial person that you could be, you would still have a problem — namely, when your knowledge or expertise is imperfect, you really don’t know it. Left to your own devices, you just don’t know it. We’re not very good at knowing what we don’t know
In logical reasoning, in parenting, in management, problem solving, the skills you use to produce the right answer are exactly the same skills you use to evaluate the answer
Surveys – Should we believe them?
There is a lot written about survey methodology as a tool in evaluation, but despite the easy and neat stats that they deliver, one should regard them with a little bit of skepticism, it seems.
Two stories to demonstrate the point:
According to a speaker on 702 talk radio I heard earlier this week, Volkskas bank still receives votes for one of the best brands in South Africa (in the Annual Markinor survey), despite the fact that it has ceased existence now for more than just a couple of years. At least in this survey, you can identify problematic answers because survey respondents had the option of giving an open-ended answer. I shudder to think what people actually do when they get one of those tick box multiple choice surveys…
In the next example, it is just so clear that one should question even the most basic assumptions people make when they complete a survey.
http://news.yahoo.com/s/afp/20080204/wl_uk_afp/britainpeoplehistoryoffbeat_080204001239
LONDON (AFP) – Britons are losing their grip on reality, according to a poll out Monday which showed that nearly a quarter think Winston Churchill was a myth while the majority reckon Sherlock Holmes was real. The survey found that 47 percent thought the 12th century English king Richard the Lionheart was a myth. And 23 percent thought World War II prime minister Churchill was made up. The same percentage thought Crimean War nurse Florence Nightingale did not actually exist.Three percent thought Charles Dickens, one of Britain’s most famous writers, is a work of fiction himself. Indian political leader Mahatma Gandhi and Battle of Waterloo victor the Duke of Wellington also appeared in the top 10 of people thought to be myths. Meanwhile, 58 percent thought Sir Arthur Conan Doyle’s fictional detective Holmes actually existed; 33 percent thought the same of W. E. Johns’ fictional pilot and adventurer Biggles.



