Le déjà-vu et le pas-encore de l’espace latent / The Already-Seen and the Not-Yet of the Latent Space

Six années séparent la critique de 2016 de l’état où se trouve aujourd’hui l’espace latent, et la tentation serait grande de lire cet intervalle comme une émancipation. On dirait volontiers que la machine s’est arrachée à la ressemblance, que le bruit gaussien a ouvert une brèche par où s’engouffre enfin l’imprévisible, que la génération a cessé de recycler le passé moyen. Cette lecture est fausse, et son erreur éclaire ce qui s’est réellement produit. Ce que les architectures d’attention et les modèles de diffusion ont accompli entre 2016 et 2022 ne dissout pas l’ontologie de la ressemblance. Ils la déplacent, la temporalisent et l’industrialisent. La mimèsis n’a pas été congédiée. Elle est devenue une géométrie, une trajectoire, une métrique inscrite en amont de toute perception singulière. Le sort de la ressemblance est plus troublant que l’affranchissement qu’on serait tenté de lui prêter, car elle ne s’est pas retirée du dispositif, elle s’y est installée comme sa loi silencieuse.

La première mutation vient du langage, par l’imposition d’un régime d’attention contextuelle. Lorsque Vaswani et ses collaborateurs publient « Attention Is All You Need » (Vaswani et al. 2017), l’ambition affichée demeure séquentielle : abandonner la récurrence pour décoder le texte en parallèle à partir de matrices de similarité. Les conséquences pour la topologie des représentations excèdent de loin ce programme. Le vecteur d’un mot cesse de subir l’inertie d’un espace global et figé ; sa position se recalcule à chaque couche, à chaque pli de l’attention, selon le voisinage qui l’entoure. L’espace vectoriel se met à réfracter là où il se contentait de cartographier. Avec l’explosion des grands modèles au seuil des années 2020 (Brown et al. 2020), cette plasticité gagne l’ensemble des régimes perceptifs. Il faut cependant nommer avec exactitude ce qui se plastifie. La ressemblance ne s’assouplit pas ; c’est la manière dont elle se calcule qui gagne en finesse contextuelle. L’attention raffine la métrique de proximité, elle en resserre le grain, elle ne la lève jamais. Un mot devient une position mobile dans un champ de forces, mais la mobilité elle-même reste réglée par la distance apprise entre les coordonnées.

Le second déplacement, plus spectaculaire, tient à l’abandon progressif des réseaux antagonistes au profit des modèles de diffusion, et c’est de lui que naît l’illusion d’affranchissement. Le réseau antagoniste projetait d’un seul geste un vecteur latent brut à travers un générateur, dans un face-à-face tendu avec un discriminateur ; la ressemblance y était instantanée, livrée d’un coup par une opération de décodage. La diffusion procède selon une économie temporelle inverse, presque liturgique, qui commence par détruire pour réapprendre à faire venir. Ho, Jain et Abbeel en donnent la formulation canonique dans « Denoising Diffusion Probabilistic Models » (Ho, Jain, and Abbeel 2020), amplifiée l’année suivante par Nichol et Dhariwal (Nichol and Dhariwal 2021) : on prend une image, on y verse par étapes un bruit gaussien jusqu’à dissolution complète de la structure, puis on entraîne un réseau, le plus souvent un U-Net doté de blocs d’attention, à prédire et retrancher ce bruit à rebours. La forme ne surgit plus d’un point ; elle décante d’un bruit blanc à travers des centaines d’itérations guidées, et chaque pas de calcul est une décision différentielle prise au sein d’un espace de probabilités.

C’est ici que la critique doit se garder d’une facilité. Il serait commode de conclure que le bruit initial, réservoir d’entropie qui n’appartient à aucun échantillon particulier du corpus, introduit une indétermination absolue, une matrice d’aléa pur d’où sortirait une altérité sans dette. La description est doublement fautive. Le bruit de départ n’appartient effectivement à aucune image d’entraînement, mais la trajectoire qui le résout appartient tout entière aux poids, donc à l’ingestion massive dont ils gardent la trace. L’indétermination du point de départ est réelle ; elle est conditionnée de part en part par la distribution apprise qui la reconduit vers une forme. Ce que la diffusion sculpte, ce n’est pas le chaos en lui-même, c’est un chaos dont chaque décantation obéit à la géométrie sédimentée du dataset. La ressemblance n’a pas disparu du dispositif ; elle a changé de temporalité. Elle s’est étalée sur la durée d’un débruitage, elle a cessé d’être donnée pour devenir processuelle, et cette processualité est précisément ce qui la rend méconnaissable à qui la cherchait sous la forme d’un décodage immédiat.

Le philosophe qui lisait dans l’espace latent de 2016 un simple redéploiement du possible bergsonien, un déjà-là masqué par un artifice de lissage, n’est donc pas pris de court par un dehors radical. Il est confronté à un régime du possible plus retors, que le vocabulaire de la probabilité laplacienne manque entièrement. Un visage produit par ces réseaux n’est pas la moyenne probabiliste des visages du corpus ; il est une entité inédite, cohérente avec le tissu du possible, dotée de traits qui n’appartiennent à aucun visage réel. La ressemblance qu’il exhibe ne signale pas une probabilité, elle signale une possibilité. Ce déplacement du probable vers le possible constitue le vrai gain conceptuel de la décennie, et il se paie d’un renversement : ce que la machine touche, ce ne sont pas des occurrences existantes qu’elle estimerait, mais des états cohérents qui n’ont jamais eu lieu et que la trajectoire de débruitage porte à l’existence sans les extraire d’aucune archive.

L’étape qui soude l’attention textuelle et la diffusion stochastique se concrétise en 2022 avec les modèles de diffusion latente, que Rombach et ses coauteurs fixent dans « High-Resolution Image Synthesis with Latent Diffusion Models » (Rombach et al. 2022). Appliquer le débruitage à des millions de pixels en haute résolution exigeait une puissance démesurée ; le groupe de Heidelberg scinde le problème. Un auto-encodeur compresse l’image dans un espace latent de dimension réduite, qui écarte le détail superflu et retient la sémantique structurelle, et le processus de diffusion opère à l’intérieur même de cet espace compressé. Pour que cette machinerie aveugle réponde à une intention, on greffe sur ses couches intermédiaires un mécanisme de cross-attention issu des transformers, alimenté par les représentations textuelles d’un modèle contrastif comme CLIP (Radford et al. 2021). Le texte cesse d’agir comme une étiquette isolée. Il devient un champ de forces qui pilote en temps réel la pente du débruitage à l’intérieur de l’espace latent intermédiaire. Écrire une invite ne sélectionne plus un point fixe dans une topologie immobile ; l’invite incline les versants d’un paysage stochastique, elle canalise le bruit vers une cuvette sémantique. Mais cette cuvette a été creusée par le corpus, et le texte qui l’oriente est un opérateur métrique : il repondère les relations de proximité déjà inscrites dans l’espace commun aux mots et aux images. Le séisme culturel de 2022 tient à ceci, et non à une libération. Le jugement de ressemblance, qui supposait autrefois un acte du regard toujours révisable, s’automatise et se fait pilotable par la langue en temps réel. La machine a internalisé sous forme de géométrie ce qui commandait jadis la production des images.

On peut mesurer là ce que devient la latence. L’image générée existait déjà à titre de possible statistique, comme pré-contenu, avant toute production ; sa coordonnée était disponible dans l’espace vectoriel bien qu’aucun support ne l’eût jamais enregistrée. Or rien, dans cet espace, ne sépare l’image apprise de l’image qui n’a pas encore eu lieu. L’image du dataset occupe une région plus dense, la région où la probabilité de la retrouver est plus haute ; l’image inédite occupe une région plus rare. L’écart entre les deux est de degré, non de statut ontologique. Aucune frontière ne partage le déjà-vu du pas-encore-advenu. Cette commensurabilité suffit à courber la flèche du temps, car la trajectoire de débruitage ne traverse à aucun moment une limite qui séparerait la mémoire de l’invention. Elle descend un gradient de densité. C’est pourquoi l’expression de fiction sans origine décrit mal ce qui advient. L’origine n’est pas absente. Elle est distribuée dans les poids et brouillée dans le temps, disséminée au point qu’aucun échantillon ne siège aux coordonnées de l’artefact produit, alors que chacune des dimensions de cet artefact procède du corpus.

Reste la question de la nouveauté, qu’il faut poser dans les termes justes. Les hybrides improbables, les monstruosités stylistiques que ces modèles prolifèrent existent bel et bien, et ils déjouent la vraisemblance statistique tout en s’y appuyant. Leur nouveauté demeure combinatoire, interne au tissu du possible. Ce qui s’automatise dans ce régime, c’est la mimèsis elle-même, et le produit porte une double reconnaissance : la trace des images d’origine, et la signature stylistique du réseau, ce flou, cette qualité de gris, ce grain de bruit qui composent une esthétique de l’erreur féconde. La disproduction tient à une ressemblance doublée d’une dissemblance, une différence à la limite de la répétition, reconnaissable sans être identique. Tant qu’elle s’accumule sans se composer, sur un espace latent commercial qui n’appartient plus à personne, cette différence reste neutre, non informative au sens de Bateson, une différence qui ne fait pas différence. Elle ne devient informative qu’à se mesurer contre la clôture d’un corpus fini, contre la finitude d’un ensemble qui a cessé de croître parce que quelqu’un ne peut plus rien y verser, et où la dissimilarité se rapporte enfin à une perte. Les deux régimes participent d’une même contrefactualité généralisée, d’états du monde qui auraient pu exister : l’un les produit parce qu’il le peut, l’autre parce qu’il le doit.

L’espace latent de 2022 n’est donc pas le grenier où gisait le double appauvri des archives, et il n’est pas davantage un plan de fictions sans origine où la machine aurait apprivoisé un chaos souverain. Il est le lieu où le jugement de ressemblance est devenu une machine, où le possible se produit comme possible, où l’invention et la mémoire partagent la même métrique et cessent pour cette raison de s’opposer. La philosophie se trouve requise de penser une mimèsis qui a quitté l’acte du regard pour devenir une propriété de l’espace, une ressemblance qui ne se discute plus et se mesure. Le disréalisme ne consiste pas à briser ces miroirs. Il consiste à mieux savoir user de leurs inévitables effets déformants, à habiter cette divergence minimale par laquelle la technique se met parfois à différer de ce pour quoi elle avait été produite.


Six years separate the 2016 critique from the state in which the latent space finds itself today, and the temptation would be great to read this interval as an emancipation. One would gladly say that the machine has torn itself away from likeness, that Gaussian noise has opened a breach through which the unpredictable finally rushes in, and that generation has stopped recycling the average past. This reading is false, and its error illuminates what actually occurred. What attention architectures and diffusion models accomplished between 2016 and 2022 does not dissolve the ontology of likeness. They displace it, temporalize it, and industrialize it. Mimesis has not been dismissed. It has become a geometry, a trajectory, a metric inscribed upstream of any singular perception. The fate of likeness is more troubling than the liberation one might be tempted to attribute to it, for it has not withdrawn from the apparatus; it has installed itself there as its silent law.

The first mutation comes from language, through the imposition of a contextual attention regime. When Vaswani and his collaborators published “Attention Is All You Need” (Vaswani et al. 2017), the stated ambition remained sequential: abandoning recurrence to decode text in parallel using similarity matrices. The consequences for the topology of representations far exceed this program. A word’s vector ceases to suffer the inertia of a global and frozen space; its position is recalculated at each layer, at each fold of attention, according to the surrounding neighborhood. The vector space begins to refract where it once merely mapped. With the explosion of large models at the threshold of the 2020s (Brown et al. 2020), this plasticity extends to all perceptual regimes. However, one must name precisely what is being plasticized. Likeness does not become more flexible; it is the way it is calculated that gains in contextual finesse. Attention refines the proximity metric, it tightens its grain, but it never lifts it. A word becomes a mobile position within a force field, but the mobility itself remains governed by the learned distance between coordinates.

The second displacement, more spectacular, stems from the gradual abandonment of generative adversarial networks in favor of diffusion models, and from it arises the illusion of liberation. The adversarial network projected a raw latent vector in a single gesture through a generator, in a tense face-off with a discriminator; likeness there was instantaneous, delivered all at once by a decoding operation. Diffusion proceeds according to an inverse temporal economy, almost liturgical, which begins by destroying in order to relearn how to bring forth. Ho, Jain, and Abbeel give its canonical formulation in “Denoising Diffusion Probabilistic Models” (Ho, Jain, and Abbeel 2020), amplified the following year by Nichol and Dhariwal (Nichol and Dhariwal 2021): an image is taken, Gaussian noise is poured into it step by step until the structure is completely dissolved, and then a network—most often a U-Net equipped with attention blocks—is trained to predict and subtract this noise in reverse. Form no longer surges from a point; it decants from white noise through hundreds of guided iterations, and each calculation step is a differential decision made within a probability space.

It is here that critique must guard against a facility. It would be convenient to conclude that the initial noise, a reservoir of entropy belonging to no particular sample of the corpus, introduces an absolute indetermination, a matrix of pure randomness from which an alterity without debt would emerge. The description is doubly flawed. The starting noise indeed belongs to no training image, but the trajectory that resolves it belongs entirely to the weights, and thus to the massive ingestion of which they keep the trace. The indetermination of the starting point is real; it is conditioned through and through by the learned distribution that leads it back toward a form. What diffusion sculpts is not chaos in itself, but a chaos each decantation of which obeys the sedimented geometry of the dataset. Likeness has not disappeared from the apparatus; it has changed temporality. It has spread out over the duration of denoising, it has ceased to be given in order to become processual, and this processuality is precisely what makes it unrecognizable to anyone seeking it in the form of immediate decoding.

The philosopher who read in the 2016 latent space a simple redeployment of the Bergsonian possible, a already-there masked by a smoothing artifice, is therefore not caught off guard by a radical outside. They are confronted with a more devious regime of the possible, which the vocabulary of Laplacian probability entirely misses. A face produced by these networks is not the probabilistic average of the faces in the corpus; it is an unprecedented entity, coherent with the fabric of the possible, endowed with traits belonging to no real face. The likeness it exhibits does not signal a probability; it signals a possibility. This displacement from the probable to the possible constitutes the true conceptual gain of the decade, and it comes at the price of a reversal: what the machine touches are not existing occurrences that it would estimate, but coherent states that have never taken place and that the denoising trajectory brings into existence without extracting them from any archive.

The step that welds textual attention and stochastic diffusion materializes in 2022 with latent diffusion models, which Rombach and his co-authors set down in “High-Resolution Image Synthesis with Latent Diffusion Models” (Rombach et al. 2022). Applying denoising to millions of pixels in high resolution required disproportionate power; the Heidelberg group splits the problem. An autoencoder compresses the image into a reduced-dimension latent space, which discards superfluous detail and retains structural semantics, and the diffusion process operates directly inside this compressed space. For this blind machinery to respond to an intention, a cross-attention mechanism from transformers is grafted onto its intermediate layers, fed by the textual representations of a contrastive model like CLIP (Radford et al. 2021). Text ceases to act as an isolated label. It becomes a force field that pilots in real time the slope of denoising within the intermediate latent space. Writing a prompt no longer selects a fixed point in an immobile topology; the prompt tilts the slopes of a stochastic landscape, it channels noise toward a semantic basin. But this basin has been hollowed out by the corpus, and the text orienting it is a metric operator: it reweights the proximity relations already inscribed in the common space of words and images. The cultural earthquake of 2022 lies in this, and not in a liberation. The judgment of likeness, which once presupposed an act of the gaze that was always revisable, is automated and becomes pilotable by language in real time. The machine has internalized in the form of geometry what once commanded the production of images.

One can measure here what latency becomes. The generated image already existed as a statistical possible, as pre-content, before any production; its coordinate was available in the vector space even though no medium had ever recorded it. Yet nothing in this space separates the learned image from the image that has not yet taken place. The dataset image occupies a denser region, the region where the probability of finding it is higher; the unprecedented image occupies a rarer region. The gap between the two is one of degree, not of ontological status. No boundary divides the already-seen from the not-yet-occurred. This commensurability is enough to bend the arrow of time, because the denoising trajectory at no point crosses a limit separating memory from invention. It descends a density gradient. This is why the expression of fiction without origin poorly describes what happens. The origin is not absent. It is distributed in the weights and blurred in time, disseminated to the point where no sample sits at the coordinates of the produced artifact, whereas each of the dimensions of this artifact proceeds from the corpus.

There remains the question of novelty, which must be posed in the right terms. The improbable hybrids, the stylistic monstrosities that these models proliferate do indeed exist, and they thwart statistical verisimilitude while relying on it. Their novelty remains combinatorial, internal to the fabric of the possible. What becomes automated in this regime is mimesis itself, and the product bears a double recognition: the trace of the original images, and the network’s stylistic signature, that blur, that quality of gray, that grain of noise that make up an aesthetic of fertile error. Disproduction stems from a likeness coupled with a dissemblance, a difference on the verge of repetition, recognizable without being identical. As long as it accumulates without composing itself, on a commercial latent space that belongs to no one anymore, this difference remains neutral, uninformative in Bateson’s sense, a difference that makes no difference. It only becomes informative when measured against the closure of a finite corpus, against the finiteness of a set that has stopped growing because someone can no longer pour anything into it, and where dissimilarity finally relates to a loss. Both regimes participate in the same generalized counterfactuality, in states of the world that could have existed: one produces them because it can, the other because it must.

The latent space of 2022 is therefore not the attic where the impoverished double of archives lay, nor is it a plane of fictions without origin where the machine would have tamed a sovereign chaos. It is the place where the judgment of likeness has become a machine, where the possible is produced as possible, where invention and memory share the same metric and for this reason cease to oppose each other. Philosophy is required to think a mimesis that has left the act of the gaze to become a property of space, a likeness that is no longer debated but measured. Disrealism does not consist in breaking these mirrors. It consists in better knowing how to use their inevitable distorting effects, in inhabiting this minimal divergence by which technique sometimes begins to differ from what it was produced for.