Latent Space Navigators

I recently read Kevin Kelly's substack essay on latent spaces, and it's all I can think about. Kelly points to the potential within LLMs that exists because they at once have an extremely saturated collection of human documented knowledge encoded into vectors within a model, and at the same time, an even more vast collection of the space between these encoded points.

The challenge here is that we are navigating between the latent spaces of an AI's training model. Even at a trillion- or trillion-and-a-half-parameter model, that doesn't contain the infinite possibility of what could be created or said or sung or drawn or whatever. Each user's intent, more likely than not, lies in between the points in the training data space or the parameters of the model.

The exciting potential of working with a LLM is navigating to new spaces between these points and finding out what lies there - the unrendered potential.

We (you, me, the model) have seen a cat. We have seen a hat. We have even seen a cat wearing a hat. But we have never seen the specific kind of cat, with the exact number of stripes and color eyes and shape of its nose and whiskers, that I am trying to imagine right now. But that exact combination exists out there, in the model, between the known knowns, in those latent spaces, and we can go find it.

When I am using a LLM I am excited by the challenge of this exploration. How can I summon the particular outcome I am aiming at into existence.

Getting that particular outcome is often very difficult. As I have experienced, and probably many of you have experienced, LLMs are unpredictable tools that need a lot of steering to get towards a predictable and particular outcome that you have in mind. And that's because I think it's difficult to push them away from those centers of gravity, those known vectors, the weights within the model. The outcomes are pulled towards them.

I imagine when I am navigating these latent spaces, it's a lot like navigating the galaxy of stars and planets and asteroids. Your ship wants to be caught in the gravity well of one of these objects and be pulled down towards that object, which then becomes the outcome. The challenge of navigating to a particular latent space is not allowing that to happen.

And it's pretty fun when you think about it that way, or at least I find that challenge very motivating. Granted, I am often doing this exploration as a hobby and out of my own curiosity, without necessary time pressures that otherwise would make this experience extremely frustrating.

But I have noticed it has gotten less frustrating over time, at least to achieve particular outcomes that are simple, like generating a particular line of copy or image, or the basic structure of a game or website. Even with particularities, I have noticed it's much easier to navigate to those latent spaces, and with those wins, I've just continued to push my ambition. What other latent spaces can we achieve and navigate to that are more complicated, that require navigating between even more gravity wells of known knowns to get to that triangulation of a new and complex point among these probabilistic stars?

And this experience of navigating to the point, between the points, getting to the specific outcome, I find to be extremely creative because you have to figure out what levers you have to pull to pilot the LLM to go somewhere that has never been before within its own parameters. It is a puzzle! You have to navigate these latent spaces, and so we become latent space navigators. That is the challenge of using these tools well, bringing our own understanding of what good looks like. This is everybody's hyping up of taste and judgment and trying to achieve it.

I think that's what we're doing when we're using LLMs. We're becoming latent space navigators.

How do I quickly take steps in a direction, assess if it's the right direction, reorient, and map where I've been and where I'm trying to go? When there is no map to that destination, we are literally explorers of this space. We don't know how to get to where we want to go. We barely know where we want to go, and to me, that's just a very, very exciting experience in the digital space. It's way, way different than all of the experiences up until this point, which have been way more predictable and deterministic, or even overly predictable. If we're thinking about the attention economy and the algorithms we get served up, there's no algorithm here. Okay, there are algorithms, but not in your interaction with the LLM in that experience, because it's a probabilistic tool. Your inputs create a spread of potential solutions that are then whittled down to one, and so this has just been a blast of exploration and creativity.

And we have all these emerging tools on how to navigate that space effectively. They're all over the place. You have:

  • prompting and context engineering
  • MCP servers with lots of documentation and tools that LLMs can use
  • iterative loops
  • coding harnesses and a wide range of other harnesses
  • And you have an increasing number of surfaces to communicate and navigate together with the LLM, such as Claude Code's playgrounds, artifacts, sketch renderings in HTML, and so on.

And I think the UI of this navigation is just in its infancy. Right now, I like to work in VS Code with my explorer open so I can see the context and manipulate it with Claude Code in terminal and context and images and HTML files pulled up in the previewer. These give me enough of a surface to do that triangulation of getting what I imagine the possible outcome I'm aiming for is in as much detail as possible out of my head and into documentation and graphs and iterations, which are all steps through that latent space, trying to prevent us from falling into an outcome too early. One of my most recent experiments with using AI creatively was building an engine that could intake an essay like this one and design a front-end rendering of the essay with interactive elements that would support communicating the points of the essay beyond what the words could achieve alone. And it was a really fun, challenging experience to build this engine. One of the insights from it was that, in order to get Claude to push further beyond generic and templated outputs, I had to slow it down. I had to slow it down by creating a flow and a system of maximum consideration:

  • Ingest the essay.
  • Read it.
  • Understand it.
  • Map out the beats.
  • Make a draft composition of where visualizations should exist.

Ultimately, what we landed on was trying to create a mind's eye for Claude: how do we get it to imagine, in the same way that I do, what possibly could come to life in the essay? And the way we did that was to force it to generate, literally in one run, 2,700 images using an index of front-end design techniques that it also had to parse through and make selections from. By doing this, it slowed Claude down. It had to actually generate all these sketches of outputs and then review them and make decisions. It couldn't just leap to a decision from text to code. And that was a navigational technique. I had to come to a realization of how we navigate between the averages, between the weights and their gravity, to get to a new triangulation where a really effectively and particularly rendered version of this essay lives.

I'm pretty happy with the results. There's definitely always room for improvement, but in terms of exploring techniques in how to push these tools and myself to navigate the latent spaces, this has been a really rewarding challenge to tackle. I think, in general, trying to get Claude to be particularly expressive or creative in a way that reflects what you have in mind is an awesome challenge. I think we are going to need to continue to develop better and better and better UI for doing that navigation together.

And so that's what I'm gonna start exploring next: what is the right UI of latent space navigation in particular around creative projects? I imagine this can be applied to any kind of effort to get an LLM to produce a particular kind of output, from business memo to mathematical equation or protein folding or what have you.