an interactive exercise
Visual Literacy in the Age of Machine Vision
Algorithms are constantly feeding us images designed to hold our attention for as long as possible, through outrage, fear, entertainment, and manipulation. That makes it more important than ever to understand what makes a human's reading of a photograph different from the way a machine registers it.
human
A human reads for meaning
When you look at a photograph, you're not decoding pixels — you're inferring a world. Who are these people, and how do they relate to each other? What just happened, or is about to? What's outside the frame? Does it make you feel something? You bring your own history to it, and no two viewers read it quite the same way. Ambiguity isn't a failure here; it's often the point.
machine
A machine reads for classification
A computer vision model never sees a "scene." It sees a grid of numbers, such as brightness and color values, converted into statistical patterns matched against categories it was trained on. It outputs whatever labels its training data taught it to associate with similar patterns. There's no ambiguity tolerance, no narrative, no nuance.
Why this gap matters
Most images made and processed today are never seen by a human eye at all.
They're captured, sorted, and acted on entirely by machines: facial recognition systems, attention trackers, industrial sensor networks, autonomous surveillance systems.
Investigative artist Trevor Paglen calls these "invisible images": images made by machines, for machines, operating on and shaping our behavior without ever passing through human perception. When a photograph is read this way, "understanding" isn't the goal at all — sorting, flagging, commodification, and targeting are.
Library scientists are warning about the implications. ACRL's own visual literacy standards now urge learners to "assess how emerging technologies such as deep fakes, facial recognition, and other applications of artificial intelligence may impact visual perception, privacy, and trust."
About this exercise: what you will experience
This interactive exercise uses a photograph of your choosing to walk through both readings. First, a human's rich, ambiguous process. Then, a machine's flat, confident-sounding labels. And finally, a chance to compare the two. For best results, set aside about 10 minutes to go through it all. There's a lot to learn here!