Skip to main content

Measuring Creativity

 

Creativity is a defining human characteristic, but it is also notoriously hard to pin down. I recently wrote about the benefits of diverse perspectives in solving real-world problems, which reminded me of the psychology literature on measures of creativity. We psychologists specialize in measuring characteristics of people that are hard to define, like intelligence or personality. But within the broad literature of psychological testing, creativity is a particularly tough nut to crack.

Dr. Teresa Amabile of the University of Iowa pointed out that all studies of creativity begin with a significant "criterion problem," in other words, how do we define the thing that we want to study? To some extent, whether or not a work is "creative" is in the eye of the beholder. In my own Master's thesis on creativity, I operationally defined creativity as "being a student in art school," which I compared to "being a business student." But are artists inherently more creative than businesspeople? A biography of Steve Jobs might suggest otherwise.

The psychological study of creativity has largely defined creativity as "divergent thinking," something that can be measured more easily than real-world creative output. One classic divergent-thinking test is Guilford's 1967 alternative uses test, which asks people to think of as many answers as possible to the question "what can you do with a brick?" (Here are some different answers). Another common divergent-thinking test is the "dots and lines" problem, one variant of which is shown below: 

Speakers and consultants like this one because it suggests the need for thinking that is literally "outside the box" (or at least outside the square). In psychological terms, solving this problem requires "breaking set," which is the inherent visual bias toward seeing the dots as a single shape with discrete boundaries. The same goes for the alternative uses test, where a brick could be a construction component, a paperweight in an office, a weapon in a fight, a heat-storage device on a cold night, or many other things depending on the context in which you imagine it. A third common creativity measure is the remote associates test, which requires identifying the common link between 3 words that at first glance seem unrelated (the New York Times daily "connections" puzzle works on this same principle). 

So we can measure divergent thinking, and it seems at least different from traditional intelligence tests. But is it really a good yardstick for creativity? Merriam-Webster has a somewhat circular definition of creativity as "the ability to create." Cambridge Dictionary says it's "the ability to produce or use original and unusual ideas." That's closer to what psychologist Robert Sternberg, another creativity researcher, says: creativity is the ability to produce something that's both "original and worthwhile." Psychologist Dean Keith Simonton argues for 3 qualities: original, useful, and also surprising. The difference between "original" and "surprising" is whether the person would have thought of it eventually -- to be truly creative, he says, a creative work needs to be something outside of the person's prior range of expertise. Things that are just original and useful could otherwise be highly skilled technical products, and not genuinely creative works. 

I'm dwelling on Simonton here because he has also been an advocate for the "genius and madness" hypothesis in which creativity is connected to mental illness or instability. The criterion of "surprise" means that a creative solution cannot have been available to the creative person's conscious (Narrative) mind before the act of creation. This requirement leads us down a rabbit-hole of free association, psychedelic experiences, and alcoholism, all methods by which people have tried to unleash their creativity. It also seems connected to the divergent thinking tests: The more unusual the answer, and the larger the number of unusual answers that can be produced, the better someone's score will be on some of the divergent-thinking tasks (the alternative uses test in particular). Many of these propositions have been identified by other psychologists as myths of creativity. In particular, a work doesn't always need to be completely new in order to be seen as creative. Simonton acknowledges this in considering scientific creativity, which he thinks is often more gradual and iterative than artistic creativity. But is that just because he's a scientist and not an artist? Even the works of Renaissance geniuses can be seen as iterative improvements over time, building new painting or sculpting techniques gradually on the models of artists who came before them.

The real measure of a creativity test, of course, should be whether it predicts creativity. In this area the science is far from impressive. Even advocates define divergent thinking tasks as “reasonably valid” measures of creativity, and often resort to redefining what “creativity” means. Lack of standardization is a problem: There are a wide variety of creativity assessments in use, with scoring criteria that are relatively fluid and have changed over time. Even with respect to a single scoring dimension like “originality,” rater instructions can vary widely from one study to another. Conceptually, success on divergent thinking tests also seems to require some level of convergent or non-creative thinking to whittle down the options, consistent with my Two-Minds argument that creativity involves a mix of Intuitive and Narrative thought. At best, tests of divergent thinking seem to predict creative results only in certain narrow domains. 

In general, divergent thinking tasks seem to have weak concurrent or predictive validity for real-world creative performance. In a study of drawings rated by art teachers, for instance, a measure of “novelty” (surprise?) was distinct from originality and fluency (i.e., usefulness/appropriateness), and more strongly associated with the judges’ rating of how “creative” a drawing was. An intervention designed to increase divergent thinking didn’t help to increase the creativity ratings. Divergent thinking is different from what we mean when we say something is "creative."

Here are four possible contributing factors to the divergent thinking tests' poor performance:

  • divergent thinking tasks focus on novelty but not appropriateness. This is exacerbated because there are no inherent goals in “think of as many uses as you can for a brick” – the task is not tethered to the real world.
  • Real-world problem solving involves much more convergent thinking to achieve an “integrated creative process” – e.g., the difference between having interesting ideas and being able to translate them into a form that’s useful.
  • Real-world problem solving requires domain-specific expertise, not just interesting ideas. Much of the real-world creative process happens in the negotiation between the ideas in your mind and the raw materials of creation.
  • The measurement scales used in common creativity inventories such as “originality,” “fluency,” and “flexibility” are not independent of one another – e.g., works judged as non-fluent are rarely judged as being original.

There are also some important confounding variables that contribute to divergent-thinking scores: Cultural context is a strong predictor of divergent-thinking test results, e.g. with typically lower scores seen in Asian cultures that value holism and balance. Divergent thinking is reliably associated with openness to experience, a personality trait that may not have anything to do with creativity. And studies have shown that just spending additional time on a divergent-thinking task reliably produces more “creative” results. Fluency predicts “creative” outputs in the first couple of minutes, then decreases in importance. Deliberate efforts to be creative matter more as time goes on. This pattern suggests that executive functioning (vs. free association) is more strongly involved in the generation of additional “novel” responses. Again, these are results that are judged as "creative" by the divergent-thinking criteria, but the process itself suggests something other than what our folk idea of "creativity" requires.

A final piece of recent evidence comes from research on artificial intelligence. AI systems can now reliably exceed human results on divergent-thinking tasks, for example by linking multiple AI models to challenge one another’s results.  Nevertheless, the outputs of LLMs that score high on divergent thinking are still considered less “creative” than those produced by skilled human creators. There's something formulaic about their responses, despite the AI outputs meeting psychologists' formal scoring criteria for "divergent" work. Something other than divergent thinking is thus involved in creativity.

Comments

Popular posts from this blog

What You Believe about Beliefs is Probably Wrong

The Stanford Encyclopedia of Philosophy  defines a "belief" as "the attitude we have, roughly, whenever we take something to be the case or regard it as true. ... Most contemporary philosophers characterize belief as a 'propositional attitude,' [where] propositions are generally taken to be whatever it is that sentences express. For example ... 'snow is white.'" Beliefs, then, (a) are expressed in language, (b) refer to some specific contents such as "snow," and (c) express some truth about those contents. The truth need not be an empirical statement about the world -- propositions such as "x is the square root of x-squared" are also beliefs under this definition even when there is no empirical referent for "x." Beliefs can be about a single thing, or about the relationships between things, in which case they might or might not be expressed as formal rules: e.g., "every bird has wings." Language, representation...

In Pursuit of Digital Badges

  What's the appeal of digital badges? I used to think that only a crazy person would be motivated to exercise just to move some electrons around on their smartphone screen. My phone would notify me that a new challenge was available each month, and I would routinely ignore it. I felt no pain or consternation about this -- in fact, I never gave it a second thought, or never notice the challenge at all. But for the past few months it seems that I'm among those who increase their activity level just because their iPhone tells them to (my latest badges shown above). Why does this seemingly stupid behavior-change strategy work, and why did it start working for me when it didn't help me before? First, it's important to know that rewards do motivate behavior. In a previous post on behaviorism I provided some background on how this works by training the Intuitive Mind. Two Minds Theory posits that the final stage before a behavior is produced by the Intuitive Mind (after proc...

Some Things That AI Probably Shouldn't Do

  I'm on record endorsing the use of AI by students to improve the quality of their writing and their thinking, but also expressing concern about the potential for autonomous AI to end civilization! So what's the deal here? Am I for AI or against it? As in many areas of life, the answer is "it depends ...". In this blog post, I will look at some things that AI probably should not  be doing for us, which might help to delineate the areas in which it can be more beneficial. Let's start with ethics. Although some techno-futurists have argued that AI will eventually be better at knowing what's good for us than we are ourselves, a recent report showed that a "robo-ethicist" using large language models (LLM) showed notable flaws in its reasoning. LLMs' ethics were consistently more influenced by utilitarian thinking (do what causes the least harm or the most benefit in this specific situation) than by reasoning from first principles (Kant's id...