Your Gaze Has a Signature

No two people take in a new scene quite the same way, a new Dartmouth study finds, using AI to map where we look. 

Walk into a crowded coffee shop, and what catches your eye as you take in the scene could say as much about you as the spirals on your fingertips or the mutations in your DNA. Your eye movements are so unique, in fact, that they can be used to identify you. That’s the surprising conclusion of a new study by Dartmouth researchers in the journal Proceedings of the National Academy of Sciences.

Psychologists have long studied where people consciously or unconsciously focus their attention as they scan a new environment. While we usually come away with a similar understanding of a place, the details of how we get there, from where we look to how long, varies from person to person.

The PNAS study examines that variation.

“From the earliest moments of taking in a new environment, we make radically different choices about what we pay attention to,” says the study’s senior author, Caroline Robertson, an associate professor in the Department of Psychological and Brain Sciences. “This work suggests that our latent conceptual priorities are embedded in the signatures of our gaze.”

A conceptual priority is a kind of personal bias that shapes what jumps out at us visually. A flag and a football, for example, look nothing alike physically, but they are connected by abstract ideas like patriotism and identity. The study suggests that we spend more time seeking out information in new environments that is conceptually rich and personally meaningful to us. Further, these priorities can be used to distinguish one person from another.

Looking at where people look

Robertson studies visual attention and herself became focused on the individual differences that kept popping up in her experiments. With then-graduate student Amanda (“AJ”) Haskins, Guarini PhD ’24, the study’s first author, and former research assistant Katherine Packard ’23, Robertson had about 60 study participants put on VR headsets and immerse themselves in a series of everyday scenes, including an auto repair shop, a public swimming pool, and an airport.

The researchers then used eye-tracking data recorded by the headsets to model each individual’s gaze pattern. They created a traditional machine-learning model to recreate where participants looked in space; a vision model to recreate what objects held their attention; and a large language model, or LLM, to look at the conceptual themes that tied these objects together. Each image processed by the language mode had captions that hinted at potential storylines. For example, one caption read, “a flag that is on the wall and could be signaling national identity,” while another read, “a missile that is for military equipment and could be part of an aeronautics display.”

Participants were free to turn their heads, move their bodies, and look where they pleased in the 16 seconds allotted for each image. When the researchers analyzed their results, they found that the vision model and the LLM could identify individuals by their unique eye movements, and that the LLM, which specifically encoded their conceptual preferences, could do this most accurately. The conceptual map, in other words, was most predictive.

Objects that didn’t look particularly alike, but were conceptually linked, could be used to tell participants apart. In an office scene, for example, writing-related items like keyboards and notepads might seem more noteworthy to one individual, than architectural elements like moldings and decorative backsplashes to another individual, revealing their unique conceptual preferences.

These individual preferences were long lasting, the researchers found. When half the group returned a week later to explore a new set of scenes, gaze models built from their earlier eye-tracking data still accurately predicted what kinds of visual features would grab their attention.

“This suggests that individual differences in gaze patterns contain stable, personality-level preferences that extend beyond testing days,” says Robertson.

Though their gaze patterns varied, the participants had three main perceptual stages. In the first two seconds of taking in a new scene, their gaze focused on spatial dimensions, like the image’s horizon and center, before shifting to prominent visual elements, and after about eight seconds, to the meanings encoded in them.

That intuitively made sense to the researchers. We typically orient ourselves in space and then check out objects and people before transitioning to an interpretive mode that attempts to understand what it all means.

The richer the conceptual information that the LLM received, the more fine-grained distinctions it could pick up on. The researchers found that longer captions containing more context, such as “a hat that is on her head and could be keeping the sun from her eyes,” seemed to elicit more distinctive responses than just “a hat that is on her head.” The more distinctive the eye patterns, the easier they were for the LLM to pick up.

What the eyes give away

The researchers point out that eye-tracking data alone may not reveal our politics or personalities. But their findings suggest that VR and AR could be more intrusive than we realize, potentially giving away more personal data to advertisers than we do now with our clicks across the web.

The study may be the first to use an LLM to model visual gaze—and it will likely not be the last. The researchers are hopeful that the novel AI methods used in the study could have clinical applications.

“Individual gaze differences aren’t random, but rather, are consistent from place to place and stable over time,” says Haskins, now a postdoc at the University of California, San Diego. “That’s important if we want to use gaze as a clinical marker of conditions like autism.”

One hallmark of autism is a reduced focus on faces, but it’s been unclear whether face avoidance is more visual than conceptual. The approach the researchers used could help to distinguish between the two.

It could also make earlier diagnosis of autism possible. Symptoms can show up as early as two years old, but currently the national average age at diagnosis is four. “The sooner you could know that a child is processing the world differently, the sooner you could augment the teaching environment,” says Robertson.

The team’s next steps include exploring whether multimodal models that track both visual and cognitive attention could improve predictions further. They also want to test whether the conceptual priorities they’ve identified vary systematically across cultures or clinical groups.

Written by

Kim Martineau