De l’économie de l’attention à l’économie de l’intention / From the Economy of Attention to the Economy of Intention

Je cherche des chaussures de course sur un moteur de recherche, et il m’en suggère aussitôt d’autres, une marque que je ne connaissais pas, un modèle plus cher, une promotion qui expire dans deux heures. Rien dans cette scène n’appartient encore à ce que je veux nommer économie de l’intention. Le moteur me montre des objets, il rivalise pour mon regard parmi d’autres objets possibles, il reste tout entier du côté du média, aussi raffiné son algorithme de suggestion soit-il, aussi personnalisé qu’il paraisse. Ce qui change de nature, c’est la scène suivante, moins spectaculaire parce qu’elle ne produit justement rien à regarder. J’écris à un agent que je dois changer mes chaussures de course, pointure et budget donnés, et il achète, sans me montrer une image, sans me proposer un choix entre plusieurs modèles, sans qu’aucun objet vienne rivaliser pour mon attention. Le colis arrive. Je n’ai fait l’expérience ni d’un choix ni même d’un regard. Ce silence médiatique est une spéculation qu’il faut interroger, non parce qu’il serait plus confortable que la suggestion insistante du moteur de recherche, mais parce qu’il déplace la capture ailleurs, dans l’acte même par lequel j’ai dit ce que je voulais plutôt que dans ce qu’on m’a donné à voir. Pour nommer ce déplacement il faut d’abord regarder ce que chacune des deux économies prend pour objet technique, ce qu’elle donne à voir ou ce qu’elle fait dire, avant de comprendre ce qu’elle capture réellement.
L’économie de l’attention s’est historiquement organisée autour d’un objet privilégié, le média. Écran, flux, fil, bandeau, notification, l’objet varie mais la fonction ne change pas, il s’agit d’un objet fait pour être regardé, un objet d’exhibition qui s’offre à un regard déjà là et qui rivalise avec d’autres objets pour occuper ce regard un peu plus longtemps. Herbert Simon avait vu dès 1971 que dans un monde saturé d’information la ressource rare devient l’attention, et tout ce qui a suivi, la critique de Tim Wu sur les marchands d’attention, celle de Shoshana Zuboff sur le capitalisme de surveillance, celle d’Yves Citton appelant à une écologie de l’attention, décrit en réalité l’histoire d’un objet, le média, perfectionné pendant vingt ans pour retenir un regard qu’il ne produit pas lui-même. Le média se donne d’abord comme objet d’exposition, il se présente comme spectacle, comme surface à parcourir, et c’est précisément cette nature intuitive, montrée, qui permet à la phénoménologie husserlienne de préciser ce que cette capture laisse intact.
Chez Husserl, l’attention n’est pas la structure fondamentale de la conscience, elle est une modalisation à l’intérieur d’un champ déjà intentionnellement organisé. Dans les Ideen, elle se décrit comme un rayon du regard, un Blickstrahl, qui part du pôle du je et se porte sur tel objet plutôt que tel autre au sein d’un horizon perceptif déjà là, avec son centre thématique et sa périphérie marginale. Le média peut être décrit, du point de vue de ses effets phénoménologiques, comme travaillant à ce niveau, il propose un objet intuitif de plus dans ce champ, il en accroît l’éclat, il capte le rayon en le détournant d’un autre objet. Mais avant que le rayon se porte quelque part il y a un champ, et ce champ, le média ne le constitue jamais, il le trouve déjà donné. On peut capturer entièrement le regard d’un sujet pendant des heures sans toucher un seul instant à ce qui fait qu’il y a, pour lui, un monde donné comme visé. L’économie du média, aussi violente soit-elle dans ses effets, reste une économie de la surface intuitive, celle-là même qui me suggérait tout à l’heure une autre marque de chaussures.
Entre le média et le prompt, il y a une phase qu’il ne faut pas sauter, celle où l’IA générative produit elle-même un média selon l’intention formulée, c’est-à-dire selon le prompt. Ce moment change le régime médiatique de l’intérieur, avant même que le prompt ne devienne l’objet propre d’une économie distincte. Le média classique se donnait comme préexistant, autonome, disponible dans un champ où il rivalisait avec d’autres objets pour retenir le regard, et je pouvais y rencontrer un donné qui résistait à mon intention, qui la décevait ou la dépassait. Le média généré n’a plus cette autonomie, mais il ne perd pas pour autant toute contingence, il n’existe qu’en fonction d’une demande qu’il vient combler, calibré sur ce qui a été énoncé sans lui être identique, et il peut encore me surprendre, résister, produire ce que je n’avais pas anticipé. Ce n’est plus un objet qui attend un regard, c’est un regard qui commande désormais son objet. Mais la source de ce qui pourrait me décevoir a changé de lieu, elle ne vient plus du monde rencontré, elle vient du système qui génère. Husserl distinguait la visée signitive et l’intuition qui vient, ou non, la remplir, l’accomplissement, Erfüllung, restant toujours suspendu à ce que le monde donnerait ou refuserait. Dans l’expérience ordinaire ce suspens est constitutif, je peux viser quelque chose par le sens et me trouver déçu, Enttäuschung, quand l’intuition ne confirme pas ce que je visais, et c’est cette déception possible, cette résistance du donné à l’intention, qui fait l’épaisseur même du réel pour un sujet, ce que Sartre et Ricoeur décrivaient comme l’affrontement du désir au monde. Le média généré ne supprime pas ce suspens, il le déplace, la déception ou la surprise possibles ne viennent plus d’un monde extérieur à l’appareil mais de l’appareil lui-même, de ses propres marges d’indétermination. L’accomplissement cesse d’être une rencontre avec le monde pour devenir une rencontre avec un système conditionné par mon intention, ce qui prépare déjà le pas suivant, celui où ce même système ne se contente plus de produire un donné à ma place mais agit lui-même à partir de mon vouloir.
Le prompt n’est pourtant pas réductible à cette fonction génératrice de médias, et c’est en le regardant en lui-même, indépendamment de ce qu’il produit, que le changement de régime achève de se dire. Le prompt n’est pas un objet qu’on regarde, c’est un acte qu’on énonce. Il appartient à un tout autre ordre d’actes intentionnels, celui que Husserl analyse dès les Recherches logiques sous le nom d’intentions signitives, ces actes de signification, Bedeutungsintentionen, qui visent un sens avant toute intuition qui viendrait le remplir. Écrire un prompt, c’est mettre en mots ce que je veux, formuler explicitement, thématiquement, réflexivement, un but encore inaccompli. C’est en apparence le degré le plus élevé d’intentionnalité d’acte qui soit, plus élevé même que le clic ou le regard, puisque le sujet ne se contente pas de laisser son attention se porter quelque part, il dit, dans sa propre langue, ce qu’il vise. On pourrait croire que le passage du média au prompt est un progrès de souveraineté, que l’usager reprend enfin la main sur ce que l’écran lui avait longtemps dérobé, qu’il énonce au lieu de subir. L’achat sans image de mes chaussures de course, tout à l’heure, avait cette allure, celle d’une libération de la suggestion. Il faut maintenant regarder ce qui se passe matériellement lorsque ce prompt entre dans la machine, pour vérifier si cette apparence résiste à l’examen ou si elle s’effondre.
Le prompt, une fois découpé en tokens, devient une séquence de vecteurs plongés dans ce qu’on appelle l’espace latent du réseau, et ces vecteurs traversent des couches que l’architecture nomme elle-même couches d’attention. Pour chaque token, le réseau calcule trois projections, une requête, une clé, une valeur, et pondère les valeurs de tous les autres tokens par un score de similarité entre requêtes et clés, normalisé de façon à former une distribution. Le résultat est une combinaison pondérée, rien de plus, une moyenne dont les poids sont appris plutôt que fixés. Il faut dire tout de suite ce que cette opération n’a pas. Elle n’a pas de pôle du je d’où partirait un rayon, pas de champ perceptif préalable dont elle graduerait la clarté, pas de sujet pour qui un token serait thématique et un autre marginal. Rien, dans ce calcul, ne ressemble à la modalisation husserlienne d’un horizon donné. Si l’on s’arrêtait là, le mot attention appliqué aux transformers ne serait qu’une homonymie commode.
Mais le mot n’est pas arrivé par hasard dans ce vocabulaire. Bahdanau, Cho et Bengio l’introduisent en 2014 pour résoudre un problème précis, le goulot d’étranglement d’un vecteur de contexte à taille fixe dans la traduction automatique, en proposant que le décodeur revienne sélectivement sur certaines parties de la phrase source plutôt que de tout comprimer d’un coup, un geste que les auteurs présentent eux-mêmes comme une manière de faire faire à la machine ce que les humains font autrement, sélectionner plutôt que tout traiter à égalité. Quarante ans plus tôt, Simon avait nommé exactement ce problème à l’échelle anthropologique, une masse d’information à traiter et une capacité de traitement rare qu’il faut allouer. L’ingénierie a repris, sans le savoir précisément en ces termes, le mot même que l’économie avait choisi pour décrire la pénurie qu’elle diagnostiquait. Ce n’est pas une preuve métaphysique, c’est un fait d’histoire des idées techniques, vérifiable, qui autorise à pousser plus loin sans forcer, à condition de séparer strictement, à chaque étape, ce qui relève du fait architectural et ce qui relève de l’interprétation qu’on en propose.
Les matrices de projection vers les requêtes, les clés et les valeurs sont apprises par descente de gradient sur un corpus qui agrège des quantités massives de productions linguistiques humaines antérieures. Mon prompt est donc traité à partir de régularités statistiques apprises sur d’immenses corpus d’énoncés humains antérieurs, sans que le rapport entre ces données d’entraînement et les poids finaux soit une simple somme ou une moyenne, l’optimisation qui les produit est trop indirecte pour cela. L’interprétation, ensuite, qu’il faut assumer comme telle : on peut soutenir que l’appareil fait revenir sur mon énoncé individuel quelque chose comme une histoire statistiquement agrégée d’énoncés antérieurs. Les poids du réseau sont mathématiquement des vecteurs, et c’est seulement en ce second sens, interprétatif et non plus descriptif, qu’on peut y voir un résidu, retravaillé plutôt que figé, d’une agrégation de volontés linguistiques passées, l’attention étant l’opération par laquelle ce résidu vient s’associer à chaque énonciation singulière au moment même où elle se formule.
La génération elle-même prolonge ce trait dans le temps. Le modèle prédit un token, l’ajoute à la séquence, puis recalcule l’attention sur la séquence étendue pour prédire le suivant, indéfiniment. Il ne faut pas dire que la machine protend, au sens phénoménologique du terme, ce serait prêter à un calcul ce qui n’appartient qu’à un vécu. Il faut dire seulement qu’elle prédit, et que cette prédiction, sans sujet, sans mémoire autre que celle, gelée, inscrite dans les poids, rejoue mécaniquement la forme vide que Husserl décrivait sous le nom de protention, cette visée ouvertement indéterminée qui se détermine pas à pas à mesure que le flux avance, sans jamais l’habiter. La protention, chez le sujet, se remplissait dans l’épaisseur d’un vécu. La prédiction, chez la machine, se recalcule sans épaisseur, de token en token, à partir d’un agrégat qui ne varie pas d’un prompt à l’autre.
Le prompt n’est pas seulement reçu par le système comme une demande à satisfaire, il est lui-même susceptible d’être prélevé à son tour comme signal, absorbé dans un corpus futur qui affinera à sa suite les poids que je viens de décrire. Il faut ici distinguer des degrés qu’on confond trop facilement sous le seul mot de capture. Capturer une intention déjà formée, c’est ce que faisait le formulaire ou le sondage, recueillir un vouloir explicite pour le traiter après coup. Prédire une intention, c’est ce que fait un moteur de recherche quand il complète ma requête avant que je l’aie terminée, anticiper un contenu probable sans encore agir dessus. Influencer une intention, c’est ce que faisait déjà la publicité ciblée, orienter un vouloir par la suggestion répétée d’un objet. Ce que l’architecture décrite plus haut permet, et qui va plus loin que ces trois degrés, c’est de participer à la formation même de l’énoncé qui porte l’intention, en complétant, en reformulant, en proposant une suite avant que le sujet ait fini de dire ce qu’il voulait. Le pas ultime, celui que je ne prétends pas démontrer ici comme un fait acquis mais comme une tendance structurelle inscrite dans l’architecture elle-même, serait que l’appareil en vienne à produire l’intention à la place du sujet plutôt que de seulement la capturer, la prédire, l’influencer ou participer à sa formation.
Des chercheurs de Cambridge, Yaqub Chaudhary et Jonnie Penn, décrivaient dès décembre 2024, dans la Harvard Data Science Review, l’émergence de ce qu’ils nomment eux-mêmes une économie de l’intention, un marché naissant pour élucider, prévoir et commercialiser les signaux d’intention humaine. Leur travail décrit un phénomène économique en formation, une industrie qui se constitue autour de la commodification du vouloir. Ce que je cherche ici vient à sa suite plutôt qu’à sa place, non pas nommer ce marché mais dire ce qu’il prend pour objet, ce que signifie phénoménologiquement le fait qu’un vouloir devienne une donnée avant même d’être devenu une décision. Du côté de l’industrie elle-même, des laboratoires stratégiques comme Outlier Ventures théorisent sans détour le même basculement, non plus comme un risque mais comme un modèle d’affaires à construire, ce qui confirme, par le simple aveu de ceux qui la construisent, la réalité économique de ce marché plutôt que sa seule critique. Les poids déjà appris agrègent le passé, le prompt que j’énonce maintenant s’expose à agréger l’avenir. Ce que le langage naturel des Recherches logiques attendait comme accomplissement futur, la confirmation par le monde, se trouve remplacé par un tout autre type de suite, l’absorption du prompt dans ce qui façonnera la prochaine version de l’appareil, la mienne ou celle d’un autre. C’est là que la délégation change de nature. On dit d’ordinaire que l’agent artificiel exécute une tâche à la place de l’humain, que c’est l’agentivité qui est déléguée. Mais l’agentivité seule laisserait le sujet intact comme auteur du vouloir, simple délégation d’exécution. Ce qui se joue avec le prompt, à des degrés qu’il faut se garder d’aplatir en un seul geste, va plus loin, vers cette participation croissante du système à la formation même de l’énoncé qui porte l’intention.
Il vaut mieux chercher ce trouble, une fois l’architecture et ses degrés posés, dans une pratique d’écriture réelle que dans un geste d’interface générique. Quand j’écris avec un logiciel de complétion comme GPT-2, mon intention ne s’impose pas au texte comme un plan préexistant qui descendrait s’incarner dans la phrase. Elle se trouble au contact du dialogue qui s’instaure entre moi et le logiciel, où chacun retient ce que l’autre vient de produire et anticipe ce qu’il va produire ensuite. J’écris un début de phrase, le logiciel en propose la suite probable, je retiens cette proposition, je la corrige ou je m’en écarte, ma correction devient à son tour la matière que le logiciel retient pour anticiper la phrase suivante. La structure husserlienne de la rétention et de la protention, que Husserl décrit comme interne à un seul flux de conscience, se trouve ici distribuée entre deux systèmes, l’un vécu, l’autre calculé, sans qu’il faille prêter au second ce qui n’appartient qu’au premier. Ce n’est pas une capture unilatérale de mon intention par la machine. C’est une intention qui ne peut plus se dire pleinement mienne ni pleinement sienne, tissée dans un aller-retour où le protendu de l’un devient la donnée retenue par l’autre.
Readonlymemories VI pousse ce trouble jusqu’à son terme. Je n’y ai pas écrit un scénario, j’y ai programmé un agent qui fait émerger une histoire sans mon intention. L’agent modifie l’axe de chaque plan d’un film que je retravaille depuis des décennies, Vertigo, puis analyse les images régénérées pour prédire ce qui va se passer dans les cinq secondes suivantes, et ces prédictions textuelles deviennent à leur tour les prompts d’une nouvelle génération d’images, dans un cycle qui ne s’arrête pas. Ma position d’artiste s’en trouve troublée d’une façon que ni le concept de délégation ni celui de capture ne suffisent à décrire, parce que je suis là aussi bien l’auteur que le lecteur de l’œuvre. Je ne reçois pas le résultat comme le produit de ma volonté. Ma volonté consiste à tenter de relire ce que l’agent produit, à me relier après coup à un résultat qui procède d’une émergence proprement étrangère à ce que j’aurais pu vouloir dire. Le vouloir ne précède plus l’œuvre comme son origine, il la suit comme une tentative de relation avec ce qu’elle est devenue sans lui.
Bernard Stiegler appelait grammatisation ce processus par lequel un flux, geste, mémoire, parole, se discrétise en traces reproductibles et industriellement exploitables. Ce que l’architecture d’attention donne à voir au niveau du calcul, ce que le roman Internes et Readonlymemories VI donnent à voir chacun à son échelle, c’est une grammatisation qui ne porte plus seulement sur la mémoire ou le geste mais sur la protention elle-même, ou plus exactement sur ce qui, dans la machine, en tient lieu sans en avoir l’épaisseur, désormais discrétisé en pondérations apprises, en prédictions textuelles échangées entre deux systèmes, ou bouclé sur lui-même dans un cycle qu’aucun sujet unique ne referme plus. Merleau-Ponty donne le mot pour préciser ce partage. Il distingue, en reprenant une intuition présente chez Husserl lui-même, l’intentionnalité d’acte, thématique, judicative, celle qui pose explicitement un objet, et l’intentionnalité opérante, fungierende Intentionalität, cette couche pré-thématique, corporelle, silencieuse, qui organise le rapport au monde avant tout acte de position. Le média capturait au niveau de l’intentionnalité d’acte perceptive. Le prompt appartient d’emblée à l’intentionnalité d’acte la plus explicite qui soit, et ce même acte thématique nourrit aussitôt la constitution d’une couche opérante partagée entre l’usager et le système, celle que l’attention du réseau recalcule silencieusement à chaque token. Dans Readonlymemories VI, cette couche opérante n’est plus la mienne seule, elle circule entre mon programme et l’émergence qu’il produit, et mon intentionnalité d’acte s’y réduit à un geste de second degré, l’intention de relire plutôt que l’intention de dire.
Il faut ajouter à ce couple phénoménologique une distinction kantienne, à condition de la poser avec précision plutôt que de céder à la formule facile qui voudrait que le média s’adresse à la perception et le prompt à la raison. Rien, chez Kant, n’autorise à faire du prompt un acte de raison pure, la plupart des prompts, achète-moi des chaussures, écris-moi ce mail, restent des impératifs hypothétiques au service d’une inclination, ce que la première Critique range précisément du côté de l’hétéronomie. La distinction qui porte vraiment quelque chose est ailleurs, entre réceptivité et spontanéité. La sensibilité, chez Kant, est le lieu du donné, Rezeptivität, le sujet y reçoit une intuition qu’il ne produit pas. L’entendement, au contraire, est le lieu de la spontanéité, Spontaneität, le sujet y synthétise activement, y forme des concepts, y juge, et c’est cette spontanéité, non la seule raison au sens étroit, que Kant tient pour la marque du sujet comme tel, ce par quoi il n’est pas simplement déterminé de l’extérieur. Le média sollicite la réceptivité, il donne à voir, le sujet y reste dans la position de recevoir un objet déjà là. Le prompt peut être décrit comme un acte où la spontanéité du sujet se manifeste explicitement, le sujet y forme, y synthétise, y met en mots un vouloir qui n’était pas donné avant cet acte. Automatiser le média, c’est intervenir sur ce que Kant tenait de toute façon pour hétéronome, la réceptivité sensible soumise à ce qui lui est donné. Automatiser le prompt, c’est intervenir sur cette spontanéité même, sur le seul acte auquel Kant réservait quelque chose comme la marque propre du sujet.
Ce que Kant réservait à la spontanéité d’un seul sujet rencontre, à l’échelle industrielle, une question de nombre plutôt que de nature. Le roman Internes et Readonlymemories VI se jouent à l’échelle d’un seul auteur en dialogue avec un seul agent qu’il a lui-même programmé ou choisi. Le régime du prompt à l’échelle industrielle change l’ordre de grandeur, l’agrégation de millions de dialogues semblables en des poids uniques qui reviennent ensuite préformer chacun d’eux. C’est ici que la structure rejoint ce que j’ai appelé ailleurs le vectofascisme. Je ne désigne pas par ce mot une équivalence entre réseau neuronal et fascisme historique, mais une homologie formelle, l’agrégation de singularités exprimées individuellement en une structure collective qui revient ensuite préorienter les expressions singulières dont elle procède. Le mécanisme fasciste de mobilisation réalise cette opération par l’affect, chaque adhésion individuelle, chaque cri, chaque geste de ralliement nourrissant un affect collectif qui revient s’imposer à chaque singularité comme évidence déjà là, comme vouloir déjà su avant que le sujet ait eu à se le demander. Le réseau de neurones la réalise à froid, sans affect, dans ses propres poids. Chaudhary et Penn nomment, à propos des mêmes modèles de langage, une manipulation sociale à l’échelle industrielle, un risque social dont la forme, l’agrégation puis le retour sur le singulier, est celle que je viens de décrire. L’expulsion de la finitude, ce refus de laisser le sujet à l’épreuve de ne pas encore savoir ce qu’il veut, n’a pas besoin ici d’un chef, d’une image, d’un corps collectif rassemblé sur une place. Elle peut aussi bien se lire dans une matrice de poids que dans une foule.
On mesure alors ce que le passage du média au prompt fait vraiment apparaître, avec le média généré comme charnière entre les deux et l’architecture d’attention comme appui matériel du dernier pas, sans qu’il faille lui demander plus qu’elle ne peut donner. L’économie de l’attention monétisait le fait que je regarde. L’économie de l’intention, celle du prompt, monétise le fait que je dise ce que je veux, engageant ma spontanéité plutôt que ma seule réceptivité. Mais ce que l’achat de mes chaussures de course, au début de ce texte, donnait déjà à sentir sans le nommer encore, c’est un régime plus radical que les deux, celui où l’agent ne se contente plus de capturer, de prédire ou d’influencer mon vouloir, ni même seulement de participer à sa formation, mais menace de rendre inutile l’acte même par lequel je saurais ce que je veux avant d’agir. Je n’ai pas eu à choisir mes chaussures. C’est cette disparition de la nécessité de savoir, plus encore que la disparition du choix, qui devrait inquiéter, et c’est elle que le curseur clignotant, à la fin de ce texte, donne à éprouver plutôt qu’à comprendre. Readonlymemories VI montre qu’il existe une issue autre que la seule dénonciation, non pas une reconquête de la souveraineté perdue, mais un déplacement du vouloir vers la relecture, une manière d’habiter la position seconde plutôt que de feindre de tenir encore la première.
Il ne reste, au terme de ce déplacement, ni consolation ni synthèse à offrir. Le curseur clignote dans le champ vide où je m’apprête à écrire ce que je veux, et déjà une suggestion grisée s’y dessine, tirée de la masse statistique de ce que des millions d’autres ont formulé avant moi, pondérée par une attention qui ne regarde rien mais qui décide, malgré tout, de ce qui compte.
I am searching for running shoes on a search engine, and it immediately suggests others to me: a brand I did not know, a more expensive model, a promotion that expires in two hours. Nothing in this scene belongs yet to what I want to name the economy of intention. The engine shows me objects; it competes for my gaze among other possible objects; it remains entirely on the side of the medium, however refined its suggestion algorithm may be, however personalized it may seem. What changes in nature is the next scene, less spectacular because it produces precisely nothing to look at. I write to an agent that I need to change my running shoes, given my size and budget, and it purchases them—without showing me an image, without offering me a choice between several models, without any object coming to compete for my attention. The package arrives. I have experienced neither a choice nor even a gaze. This media silence is a speculation that must be questioned, not because it might be more comfortable than the insistent suggestion of the search engine, but because it shifts capture elsewhere: into the very act by which I said what I wanted, rather than into what I was given to see. To name this shift, we must first look at what each of the two economies takes as its technical object—what it gives to be seen or what it makes one say—before understanding what it truly captures.
The economy of attention has historically been organized around a privileged object: the medium. Screen, feed, stream, banner, notification—the object varies, but the function does not change. It is an object made to be looked at, an object of exhibition offered to a gaze that is already there, competing with other objects to occupy that gaze a little longer. Herbert Simon saw as early as 1971 that in a world saturated with information, the scarce resource becomes attention; everything that followed—Tim Wu’s critique of attention merchants, Shoshana Zuboff’s on surveillance capitalism, Yves Citton’s call for an ecology of attention—actually describes the history of an object, the medium, perfected over twenty years to retain a gaze that it does not produce itself. The medium offers itself first as an object of exposition; it presents itself as a spectacle, as a surface to be traversed, and it is precisely this intuitive, displayed nature that allows Husserlian phenomenology to clarify what this capture leaves intact.
In Husserl, attention is not the fundamental structure of consciousness; it is a modalization within a field that is already intentionally organized. In the Ideas, it is described as a ray of the gaze, a Blickstrahl, which originates from the pole of the “I” and directs itself toward one particular object rather than another within a perceptual horizon already present, with its thematic center and marginal periphery. From the standpoint of its phenomenological effects, the medium can be described as operating at this level: it proposes one more intuitive object in this field, increases its brilliance, and captures the ray by diverting it from another object. But before the ray directs itself anywhere, there is a field, and this field is never constituted by the medium—the medium finds it already given. One can entirely capture the gaze of a subject for hours without touching for a single moment that which makes there be, for them, a world given as targeted. The economy of the medium, however violent its effects may be, remains an economy of the intuitive surface—the very same surface that, a moment ago, suggested another brand of shoes to me.
Between the medium and the prompt, there is a phase that must not be skipped: the phase where generative AI itself produces a medium according to the formulated intention—that is to say, according to the prompt. This moment changes the media regime from within, even before the prompt becomes the proper object of a distinct economy. The classic medium presented itself as pre-existing, autonomous, and available in a field where it competed with other objects to hold the gaze, and in it I could encounter a given that resisted my intention, disappointing or exceeding it. The generated medium no longer possesses this autonomy, yet it does not lose all contingency for all that; it exists only as a function of a demand that it comes to fulfill, calibrated to what has been stated without being identical to it, and it can still surprise me, resist, or produce what I had not anticipated. It is no longer an object waiting for a gaze; it is a gaze that henceforth commands its object. But the source of what might disappoint me has shifted location: it no longer comes from the world encountered, but from the system that generates. Husserl distinguished between the signitive intention and the intuition that comes—or fails—to fulfill it; fulfillment, Erfüllung, always remaining suspended on what the world would grant or refuse. In ordinary experience, this suspense is constitutive: I can target something through meaning and find myself disappointed, Enttäuschung, when intuition fails to confirm what I was aiming at; and it is this possible disappointment, this resistance of the given to intention, that constitutes the very thickness of the real for a subject—what Sartre and Ricœur described as the confrontation of desire with the world. The generated medium does not eliminate this suspense; it displaces it. The possible disappointment or surprise no longer comes from a world external to the apparatus, but from the apparatus itself, from its own margins of indetermination. Fulfillment ceases to be an encounter with the world and becomes an encounter with a system conditioned by my intention—which already prepares the next step: the step where this same system no longer contents itself with producing a given in my place, but acts itself on the basis of my willing.
Yet the prompt cannot be reduced to this media-generating function, and it is by looking at it in itself, independently of what it produces, that the change of regime finishes stating itself. The prompt is not an object one looks at; it is an act one utters. It belongs to an entirely different order of intentional acts—that which Husserl analyzes as early as the Logical Investigations under the name of signitive intentions: those acts of meaning, Bedeutungsintentionen, which target a sense prior to any intuition that might fulfill it. To write a prompt is to put into words what I want—to formulate explicitly, thematically, and reflexively an as-yet unfulfilled goal. It is apparently the highest degree of act-intentionality possible, higher even than the click or the gaze, since the subject does not merely allow their attention to be drawn somewhere; they express, in their own language, what they are aiming at. One might believe that the transition from the medium to the prompt is a progress in sovereignty—that the user finally reclaims control over what the screen had long withheld from them, uttering instead of enduring. The image-free purchase of my running shoes earlier had this appearance—that of a liberation from suggestion. We must now look at what materially happens when this prompt enters the machine, to verify whether this appearance stands up to scrutiny or collapses.
Once tokenized, the prompt becomes a sequence of vectors embedded in what is called the latent space of the network, and these vectors pass through layers that the architecture itself names attention layers. For each token, the network calculates three projections—a query, a key, and a value—and weights the values of all other tokens by a similarity score between queries and keys, normalized to form a distribution. The result is a weighted combination, nothing more—an average whose weights are learned rather than fixed. We must state immediately what this operation lacks. It has no pole of the “I” from which a ray would originate, no prior perceptual field whose clarity it would graduate, no subject for whom one token would be thematic and another marginal. Nothing in this calculation resembles the Husserlian modalization of a given horizon. If one were to stop there, the word “attention” applied to transformers would be nothing more than a convenient homonymy. But the word did not enter this vocabulary by accident. Bahdanau, Cho, and Bengio introduced it in 2014 to solve a precise problem: the bottleneck of a fixed-size context vector in machine translation, by proposing that the decoder selectively return to certain parts of the source sentence rather than compressing everything at once—a gesture the authors themselves present as a way of having the machine do what humans do differently: select rather than process everything equally. Forty years earlier, Simon had named precisely this problem on an anthropological scale: a mass of information to process and a scarce processing capacity that must be allocated. Engineering adopted, without necessarily knowing it in those precise terms, the very word that economics had chosen to describe the scarcity it diagnosed. This is not a metaphysical proof; it is a verifiable fact in the history of technical ideas, which allows us to push further without forcing, provided we strictly separate, at each step, what pertains to architectural fact and what pertains to the interpretation proposed for it.
The projection matrices toward queries, keys, and values are learned through gradient descent on a corpus aggregating massive quantities of prior human linguistic productions. My prompt is therefore processed on the basis of statistical regularities learned from immense corpora of prior human utterances, without the relationship between these training data and the final weights being a simple sum or average—the optimization that produces them is too indirect for that. Next comes the interpretation, which must be owned as such: one can argue that the apparatus brings back to bear on my individual utterance something like a statistically aggregated history of prior utterances. The weights of the network are mathematically vectors, and it is only in this second sense—interpretive rather than descriptive—that one can see in them a residue, reworked rather than frozen, of an aggregation of past linguistic wills, attention being the operation through which this residue comes to associate itself with each singular utterance at the very moment it is formulated.
Generation itself extends this trait in time. The model predicts a token, adds it to the sequence, and then recalculates attention over the extended sequence to predict the next one, indefinitely. We must not say that the machine protends, in the phenomenological sense of the term; that would be attributing to a calculation what belongs only to lived experience. We must only say that it predicts, and that this prediction—without a subject, without memory other than the frozen memory inscribed in the weights—mechanically plays out the empty form that Husserl described under the name of protention: that openly indeterminate targeting that determines itself step by step as the flux advances, without ever inhabiting it. Protention, in the subject, was fulfilled within the thickness of a lived experience. Prediction, in the machine, recalculates itself without thickness, token by token, from an aggregate that does not vary from one prompt to another.
The prompt is not only received by the system as a demand to be satisfied; it is itself susceptible to being collected in turn as a signal, absorbed into a future corpus that will subsequently refine the weights I have just described. Here we must distinguish degrees that are too easily conflated under the single word “capture.” To capture an already formed intention is what a form or a survey did: collecting an explicit willing to process it after the fact. To predict an intention is what a search engine does when it completes my query before I have finished it: anticipating a probable content without yet acting upon it. To influence an intention is what targeted advertising already did: orienting a willing through the repeated suggestion of an object. What the architecture described above allows—and what goes further than these three degrees—is to participate in the very formation of the utterance that carries the intention, by completing, reformulating, or proposing a continuation before the subject has finished saying what they wanted. The ultimate step—one that I do not claim to demonstrate here as an established fact, but as a structural tendency inscribed within the architecture itself—would be for the apparatus to come to produce the intention in place of the subject, rather than merely capturing, predicting, influencing, or participating in its formation.
Researchers from Cambridge, Yaqub Chaudhary and Jonnie Penn, described as early as December 2024, in the Harvard Data Science Review, the emergence of what they themselves call an economy of intention: a nascent market for elucidating, forecasting, and commercializing signals of human intention. Their work describes an economic phenomenon in formation, an industry being constituted around the commodification of willing. What I am seeking here follows upon their work rather than replacing it—not to name this market, but to state what it takes as its object: what it means phenomenologically that a willing becomes data even before becoming a decision. On the side of the industry itself, strategic laboratories such as Outlier Ventures straightforwardly theorize the same shift, no longer as a risk, but as a business model to be built—which confirms, through the simple admission of those building it, the economic reality of this market rather than its mere critique. The weights already learned aggregate the past; the prompt I utter now exposes itself to aggregating the future. What the natural language of the Logical Investigations awaited as future fulfillment—confirmation by the world—finds itself replaced by an entirely different type of continuation: the absorption of the prompt into that which will shape the next version of the apparatus, mine or another’s. It is here that delegation changes in nature. It is usually said that the artificial agent executes a task in place of the human—that agency is what is delegated. But agency alone would leave the subject intact as the author of the willing: a simple delegation of execution. What is at stake with the prompt—at degrees that must be guarded against collapsing into a single gesture—goes further, toward this increasing participation of the system in the very formation of the utterance that carries the intention.
It is better to seek this disturbance, once the architecture and its degrees have been laid out, within a real writing practice rather than in a generic interface gesture. When I write with completion software such as GPT-2, my intention does not impose itself upon the text like a pre-existing plan descending to embody itself in the sentence. It becomes disturbed through contact with the dialogue established between myself and the software, where each retains what the other has just produced and anticipates what it will produce next. I write the beginning of a sentence; the software proposes its probable continuation; I retain this proposition, correct it, or depart from it; my correction in turn becomes the material that the software retains to anticipate the following sentence. The Husserlian structure of retention and protention, which Husserl describes as internal to a single stream of consciousness, finds itself here distributed between two systems—one lived, the other calculated—without any need to attribute to the second what belongs only to the first. This is not a unilateral capture of my intention by the machine. It is an intention that can no longer fully be called mine nor fully its own, woven in a back-and-forth where the protended of the one becomes the retained data of the other.
Readonlymemories VI pushes this disturbance to its absolute conclusion. In it, I did not write a script; I programmed an agent that causes a story to emerge without my intention. The agent modifies the axis of each shot of a film I have been reworking for decades, Vertigo, then analyzes the regenerated images to predict what will happen in the next five seconds, and these textual predictions in turn become the prompts for a new generation of images, in an unending cycle. My position as an artist is thereby disturbed in a way that neither the concept of delegation nor that of capture suffices to describe, because there I am as much the author as the reader of the work. I do not receive the result as the product of my will. My will consists in attempting to reread what the agent produces—to connect myself after the fact to a result arising from an emergence properly foreign to what I might have wanted to say. Willing no longer precedes the work as its origin; it follows it as an attempt at entering into relation with what it has become without it.
Bernard Stiegler called grammatization that process by which a flux—gesture, memory, speech—is discretized into reproducible and industrially exploitable traces. What the attention architecture reveals at the level of calculation, and what the novel Internes and Readonlymemories VI each reveal at their own scale, is a grammatization that no longer bears merely on memory or gesture, but on protention itself—or more precisely on that which, in the machine, stands in its place without possessing its thickness—now discretized into learned weightings, into textual predictions exchanged between two systems, or looped back upon itself in a cycle that no single subject completes any longer. Merleau-Ponty provides the word to clarify this division. Taking up an intuition present in Husserl himself, he distinguishes between act-intentionality—thematic, judicative, that which explicitly posits an object—and operative intentionality, fungierende Intentionalität: that pre-thematic, bodily, silent layer that organizes the relationship to the world prior to any act of positing. The medium captured at the level of perceptual act-intentionality. The prompt belongs from the outset to the most explicit act-intentionality there is, and this very thematic act immediately feeds the constitution of a shared operative layer between the user and the system—that which the network’s attention silently recalculates at every token. In Readonlymemories VI, this operative layer is no longer mine alone; it circulates between my program and the emergence it produces, and my act-intentionality is reduced to a second-degree gesture: the intention to reread rather than the intention to speak.
To this phenomenological pair, we must add a Kantian distinction, provided it is stated with precision rather than yielding to the easy formula that would have the medium address perception and the prompt address reason. Nothing in Kant justifies turning the prompt into an act of pure reason: most prompts (“buy me shoes,” “write me this email”) remain hypothetical imperatives serving an inclination—precisely what the first Critique places on the side of heteronomy. The distinction that truly carries weight lies elsewhere, between receptivity and spontaneity. Sensibility, in Kant, is the locus of the given, Rezeptivität; in it, the subject receives an intuition they do not produce. Understanding, on the contrary, is the locus of spontaneity, Spontaneität; in it, the subject actively synthesizes, forms concepts, judges, and it is this spontaneity—not mere reason in the narrow sense—that Kant holds to be the proper mark of the subject as such, that by which they are not simply determined from the outside. The medium solicits receptivity; it gives to be seen; in it, the subject remains in the position of receiving an object already there. The prompt can be described as an act in which the subject’s spontaneity explicitly manifests itself: in it, the subject forms, synthesizes, and puts into words a willing that was not given prior to this act. Automating the medium means intervening in what Kant considered heteronomous in any case: sensible receptivity subjected to what is given to it. Automating the prompt means intervening in this spontaneity itself—in the sole act to which Kant reserved something like the proper mark of the subject.
What Kant reserved for the spontaneity of a single subject encounters, on an industrial scale, a question of number rather than of nature. The novel Internes and Readonlymemories VI play out at the scale of a single author in dialogue with a single agent they have themselves programmed or chosen. The regime of the prompt on an industrial scale changes the order of magnitude: the aggregation of millions of similar dialogues into unique weights that subsequently return to preform each of them. It is here that the structure joins what I have elsewhere called vectofascism. By this word, I do not designate an equivalence between neural networks and historical fascism, but a formal homology: the aggregation of individually expressed singularites into a collective structure that subsequently returns to pre-orient the singular expressions from which it proceeds. The fascist mechanism of mobilization performs this operation through affect: each individual adherence, each shout, each gesture of rallying feeds a collective affect that returns to impose itself upon each singularity as an already-present evidence, as a willing already known before the subject even had to ask themselves about it. The neural network performs it cold, without affect, within its own weights. Chaudhary and Penn name, regarding these same language models, an industrial-scale social manipulation—a social risk whose form (aggregation, then return upon the singular) is the very one I have just described. The expulsion of finitude—this refusal to leave the subject to the trial of not yet knowing what they want—needs no leader here, no image, no collective body assembled in a square. It can be read just as clearly in a weight matrix as in a crowd.
We can then measure what the transition from the medium to the prompt truly brings to light, with the generated medium serving as the hinge between the two and the attention architecture as the material support for the final step, without demanding of it more than it can give. The economy of attention monetized the fact that I look. The economy of intention—that of the prompt—monetizes the fact that I say what I want, engaging my spontaneity rather than my receptivity alone. But what the purchase of my running shoes at the beginning of this text already allowed us to sense without yet naming it, is a regime more radical than both: one where the agent no longer contents itself with capturing, predicting, or influencing my willing, nor even merely participating in its formation, but threatens to render useless the very act by which I would know what I want before acting. I did not have to choose my shoes. It is this disappearance of the necessity of knowing, even more than the disappearance of choice, that ought to cause concern; and it is this that the blinking cursor at the end of this text allows us to experience rather than merely understand. Readonlymemories VI shows that an outcome other than mere denunciation exists: not a reconquest of lost sovereignty, but a displacement of willing toward rereading—a way of inhabiting the secondary position rather than feigning to still hold the first.
At the end of this shift, there remains neither consolation nor synthesis to offer. The cursor blinks in the empty field where I am preparing to write what I want, and already a grayed suggestion takes shape there, drawn from the statistical mass of what millions of others have formulated before me, weighted by an attention that looks at nothing, but nevertheless decides what matters.