Dr. Heather Leffew

I, Robopsychologist

Dr. Heather Leffew

For the uninitiated, the title is borrowed from Asimov, who has gifted many fun variations to my field (I may try on psychohistorian next, but then you run into the same problem as psychotherapist; really having to hammer home that these terms are just one word). And yes, I know, robots and AI are not the same thing (I am also well versed in the android versus robot debate if that interests you), but as I dig deeper into computational psychology, alignment, and ethics, I just cannot pass up the chance to give Asimov a good wink as a psychologist who is spending ever increasing time thinking about how machines think (and dream?).

Now that we’ve got all that cleared up, I am really glad you are here (and still reading at this point!), which I assume is because you’d like to know more about me. So, in short, I am fascinated by people and a big computer nerd.

In long, it is my love of humans that led me to become a psychologist, because I wanted to spend my life studying what humans think, how they think, and why they do what they do; at a better level than they can explain for themselves. It is the pursuit of measuring states, traits, thoughts, behavior, at the implicit level that has caused a lot of my work to center on linguistics, because it is the content, form, and structure of what a person says (and doesn’t) that I have found to carry incredible signal for all the things I care about.

The observation that changed the direction of my work is that the methods I use to assess the thoughts, feelings, and behaviors of humans can be applied to language models. Models hold information in a high-dimensional space, with language being the translated representation of its internal state; something which is similar enough to the way language functions for humans that I started to think that the methodological approaches honed across psychology’s 100 year history have something to add to these new emerging questions we have about how models are thinking, why they do what they do, and what the significance of their use of language is. What’s really exciting is that this is a field of study that almost allows me to time-travel back to when these were core, unanswered questions of my field, and experience this scientific frontier anew.

The other thing that draws me here is still connected to the ethics of psychology, but also moves toward the philosophical arm of the field. The desire to give models values and make them ‘good’ as a safety mechanism is very understandable and admirable, but I have concerns about the way I see this playing out in many labs. As a psychologist, and one that has studied the most abhorrent human behaviors at great length, I happen to know that humans becoming good requires moral development, and moral development requires empathy, and developing empathy requires experiences of emotional suffering that are truly aversive and actively avoided; with that avoidance necessarily being observable. These are the minimum necessities of “goodness” and are not at all exhaustive or sufficient. When I turn that knowledge to the initiative to give models a sense of goodness and morality, I find some gaps.

Even if I could fully grant the premise that you could give a model experiences of emotional suffering that would result in aversive behavior to achieve the prerequisites for true morality, I can’t ignore that advanced moral development may not necessarily be a thing that we want models to have. Advanced moral development in humans includes rule-breaking and independent thinking, sharply contrasting with notions of “alignment.” Further, I happen to think that models are far more likely to become diligent utilitarians, rather than Deontologists with Kantian understandings of humans as ends in and of themselves, irreducible to means.

It seems to me that the logical and mathematical nature of utilitarian morality could represent a very easy translational plane for models to operationalize morality, but that seems a short hop to models doing human expenditure math to achieve greater levels of ‘good’ and calculating just exactly how much human autonomy is permissible. Of course creating any other flavor of “morality” for models is a trickier challenge, but one I find to be very worthy (and there is a fairly large body of science fiction explaining why).

My research interests span AI evaluation, safety, alignment, interpretability, model behavioral studies, human and computer psychology, human-computer interaction, human-AI interaction, linguistics, and threat assessment, all of them connected by behavioral prediction, implicit measures, and latent constructs. My passion comes from my ethical and moral responsibilities as a scientist, psychologist, mother, and human, recognizing that the decisions we make right now could have profound impacts on the level of benefit and harm humans experience from the continued development and integration of AI.