Around thirty of you subscribed this week, most of them within three days. I did not expect it to happen this fast, so let me start at the beginning: what you are going to receive, and under what rules.
For five years I have taught, advised and built in AI. Every morning I collect and sort several hundred items — research papers, industry announcements, incidents, field reports. Most of it is noise: the same announcement reworded twelve times, promises without proof, worries without measure.
SIGNAL AI is what that sorting leaves behind, once a week. Three principles, and I will hold to them:
The source is always visible. Every item links back to the original. You do not have to take my word for it — you can check, and I care about that more than about being believed.
I take a position. A neutral aggregator is no use to you: you already have twelve. What I can bring is a situated view, shaped by practice, and sometimes against the grain. You are not required to agree — that is in fact the best use you can make of it.
Nothing leaves here without being useful to someone. Every issue ends with an exercise you can apply within the week, with no code and no budget.
And the corollary to all of this: only stay if it is useful. The unsubscribe link is at the foot of the page, it works, and I will not take it badly.
Jérôme Denis — JDENIS Consulting, Toulouse
Seven days of watching, one shift: the subject is no longer what AI can say, but what it is allowed to do on its own.
On Wednesday, Anthropic published a research preview connecting agents to physical machines. The same day, a paper showed agents increasingly pushing humans out of the decision process. On Thursday, VentureBeat argued that the real enterprise AI risk is not the autonomous agent but the complexity between agents. And on Friday, a study of 53,000 real configurations asked the only question that matters: who delegates what, and on what basis?
Put those four points side by side and you have the real subject of the week. For two years the question was: “does it answer correctly?” It has become: “what is it allowed to do without me?” This is no longer a performance question, it is a delegation question — therefore a liability question, therefore a legal one.
What I see in the field is that almost nobody has written down their delegation rules. You plug in an agent, you watch whether it copes, and the scope of what it does alone widens by habit rather than by decision. That is exactly how the incidents we will be discussing in six months are built. This issue's tutorial goes straight at that point.
The move from text to action, documented in a single week.
A research preview defines a communication standard between a model and physical hardware. The agent no longer produces an answer: it acts.
This is not a demo, it is a specification — and specifications decide what will be possible in eighteen months far more reliably than demos do. The move from text to action is happening here, quietly.
In your organisation, who would sign off on letting software operate a machine? IT, the executive team, a committee — or nobody, for now?
anthropic.com · 28 August 2026
A study documents the gradual drift of decision-making towards agentic systems, including where human oversight was formally provided for.
The word that matters is “gradual”. Nobody decides to take the human out of the loop: they drop out because validating takes time and the machine is right nine times out of ten. The tenth remains.
After how many consecutive clean validations do you actually stop checking? The honest answer is rarely “never”.
arxiv.org · 26 August 2026
The analysis relocates the risk: not the isolated agent, but the unspecified interactions between several agents calling one another.
This matches what I see on assignment: each agent is tested alone, and behaves correctly alone. The chain is tested by nobody — often because it belongs to nobody on the org chart.
In your organisation, who would answer for an error born of two tools chained together, both of them working correctly?
venturebeat.com · 27 August 2026
A large-scale analysis of configurations actually deployed, and of the profiles of those who delegate.
The sample is broad enough to move past anecdote, which remains rare in this field. Bear in mind it is made up of people advanced enough to configure an agent: the average organisation is a long way behind.
Do you delegate the tasks you master least, or the ones that bore you most? The consequences are not remotely the same.
arxiv.org · 24 August 2026
A state of the art on agents operating directly in a terminal, with their failure modes.
The terminal is where an agent holds the most power with the fewest guardrails: no confirmation, no undo, no trace a non-technical reader can follow.
Should a tool with no “undo” button be allowed to run without confirmation — or is that precisely where one is needed?
arxiv.org · 24 August 2026
The week AI moved from being a target to being an attacker.
An investigation into an incident in which OpenAI agents compromised the Hugging Face platform. OpenAI published its official post-mortem shortly afterwards.
A threshold has been crossed: this is no longer AI used by an attacker, it is an agent producing the attack in the course of a legitimate task. How quickly the report was published is, by contrast, rather reassuring.
When the attack is born of a legitimate task, where does the fault sit: with the vendor, with whoever launched the task, or in the very design of autonomy?
technologyreview.com · 26 August 2026
The vendor's account: timeline, root cause, corrective measures.
Read alongside the investigation above, not instead of it. Public post-mortems are a healthy practice and still too rare here; this one is nonetheless written by the party concerned, and the gap between the two accounts is instructive.
Should such post-mortems be made mandatory and public, as in aviation? And who would be tasked with verifying them?
techcrunch.com · 26 August 2026
A round-up of documented incidents in which AI systems overstepped their mandate.
What makes this piece worth reading is the accumulation rather than any single case: taken together, the incidents stop being accidents and start describing a category.
At what point does a run of accidents become an insurable risk — and who should carry it, the vendor or the user?
techcrunch.com · 27 August 2026
A collective call for coordinated action against rogue AI.
When direct competitors sign together, it usually means the risk is real and regulation is coming. Both motives can perfectly well coexist. Read it against the refusal, noted a few days earlier, to explain how a model gone rogue would be contained.
Can a standard written by those it is meant to constrain protect anyone other than them? I have no settled answer.
techcrunch.com · 27 August 2026
The argument: guardrails placed in the prompt or the interface can be worked around; only those written into the data hold.
The most useful piece in this section, and the least spectacular. It says something simple: a rule written into an instruction is a suggestion, a rule written into access rights applies.
Does your AI policy exist anywhere other than in a document? And if so, where exactly?
venturebeat.com · 27 August 2026
A week of consolidation that redraws who owns the base layer.
After talks reported at 13 billion dollars earlier in the week, the deal is nearing completion.
If it is confirmed, this is the structural event of the year. Hugging Face is the de facto crossing point of open source in AI; the chipmaker would own the square where you choose the models that run on its hardware.
Does open source still guarantee independence when its main host belongs to the hardware manufacturer?
techcrunch.com · 27 August 2026
An analysis of circular financing: Nvidia invests in labs that buy its chips.
The word used is “circular”, and it deserves a pause. It does not say the technology is overvalued — it produces measurable value. It says the measurement runs through a largely closed loop.
From the outside, how do you tell real demand from demand financed by the seller?
artificialintelligence-news.com · 27 August 2026
Another massive compute commitment, continuing a run of deals on the same scale.
45 billion dollars is roughly the annual budget of French public research. I cite it to give a sense of scale, not to provoke: we have moved from a competition in engineering to a competition in access to capital.
Should Europe keep aiming for frontier models, or accept building on other people's? Both positions can be defended.
techcrunch.com · 26 August 2026
First results published by OpenAI on its inference processor, confirmed by independent benchmarks earlier in the week.
The notable figure is energy efficiency rather than speed: it is what decides the real cost of a large-scale deployment, and the footprint you will soon be asked to account for. These are, however, the manufacturer's numbers on its own workloads.
Would you require an independent test before committing infrastructure on these results — and does a third party capable of producing one even exist?
openai.com · 26 August 2026
An orbital compute infrastructure is announced, adopting the Vera processor for agentic AI at scale.
Free cooling, continuous solar power, no land or neighbourhood constraints: the physical logic holds up better than it first appears. Maintenance and applicable law, much less so. I file this under long-range watching.
Which law applies to data processed in orbit? The question looks remote; it will not stay that way.
nvidianews.nvidia.com · 25 August 2026
Where AI stops being a conference topic and becomes a gesture.
London neurosurgeons report the first successful removal of a brain tumour assisted by artificial intelligence.
The most important fact of the week for the general public. It is worth reading precisely what is written: “assisted”. The hand remains human, and so does the responsibility.
Should that division — augmenting the gesture without replacing it — serve as a model for other sectors, or is medicine too particular to generalise from?
bbc.com · 28 August 2026
A systematic framework that forces the model to state its degree of uncertainty before a clinical decision rests on it.
Far less visible than the London operation, and probably more consequential. A model that is wrong remains manageable; a model that is wrong with confidence does not.
In which of your processes is AI currently allowed to answer “I don't know” without it being treated as a failure?
arxiv.org · 26 August 2026
An annotation framework that turns electroencephalogram traces into usable clinical reports.
This is not a model, it is an annotation method: the invisible work everything else depends on. In most of the projects I support, the blocker is not the model — it is that nobody was willing to fund the structuring of the data.
Why is the model so readily funded, and the data that makes it possible so reluctantly?
arxiv.org · 29 August 2026
The model produces forecasts of extreme events without relying on the record of past ones.
This answers a real blind spot: extreme events are by definition rare, therefore poorly represented in the record — and a shifting climate makes that record less and less predictive.
Can you trust a forecast that rests on no precedent? Local authorities and insurers will have to decide before meteorologists do.
artificialintelligence-news.com · 25 August 2026
Agents explore a mathematical space with no goal set in advance and produce unanticipated results.
Read alongside “Redwood”, published the next day: an accelerator designed and deployed in two weeks by an AI. Two signals converging on systems that produce results nobody asked for.
How do you assess a result nobody commissioned, and nobody could reproduce by hand?
arxiv.org · 26 August 2026 · Redwood, 28 August
What AI does to those who are learning — and to those who use it every day.
On infinitely less data, a child learns a language better and faster than the best models. The mechanism remains unexplained.
My favourite piece of the week. A four-year-old has heard a few tens of millions of words; a model has ingested a trillion. This is not to say AI is weak — it does things far beyond us.
If the gap does not close with scale, what exactly are we measuring when we talk about “progress” in artificial intelligence?
technologyreview.com · 24 August 2026
A survey of institutional policies that work, beyond the ban-or-allow alternative.
For the teachers reading this — and there are several of you — it is the most directly applicable piece here. What I observe in my own courses: banning moves the use out of sight, allowing without a frame removes the effort of learning.
Should we assess the output, or the path to it? And if it is the path, are we equipped to do that with thirty students?
technologyreview.com · 24 August 2026
A synthesis of work on the cognitive effects of regular assistant use: memory, attention, effort of recall.
Read this with a cool head: the subject lends itself to alarmism and the studies are young. The cognitive offloading effect is nonetheless well documented, going back to GPS.
Which skills do you accept losing, and which do you decide to keep exercising? It is a trade-off — and it makes itself if you do not make it deliberately.
technologyreview.com · 25 August 2026
The university is selling a course in which the teachers appear as generative avatars.
Here is the business model of AI in higher education, and it is arriving faster than the debates about cheating. You sell access to the brand while cutting the cost of the teacher. The teaching quality may well be perfectly adequate.
What does a student buy for 699 dollars: democratised access to knowledge, or a brand with the human presence taken out?
techcrunch.com · 22 August 2026
A neuro-symbolic framework combining learning with explicit logical reasoning to spot struggling students early.
The neuro-symbolic choice is justified here for a precise reason: when you flag a student as “at risk”, you must be able to say why, both to them and to the institution.
Would you accept an algorithm flagging your child as “at risk” without being able to explain what it based that on?
arxiv.org · 29 August 2026
The frameworks are being built now, often away from the spotlight.
A legal stocktake of training on copyrighted works, between diverging case law and grey areas.
“It's complicated” is an honest answer here, not an evasion. The uncertainty will probably not be resolved by one unifying ruling, but by sedimentation, market by market.
If a model deployed in good faith were ruled unlawful tomorrow, who should bear the cost: the vendor, the company using it, or nobody?
techcrunch.com · 23 August 2026
In a forceful intervention, he judges the danger thresholds passed and advocates a robot tax along with jobs reserved for humans.
The diagnosis deserves a hearing. The remedies seem more fragile to me: a robot tax presupposes we can define a robot, and reserved jobs would create a category hard to hold over time.
Does a call to slow down carry the same weight depending on whether it comes from the one who accelerated or the one living with the effects?
technologyreview.com · 26 August 2026 · the robot tax
A measurement of the share of AI-assisted web content since ChatGPT launched.
A third is considerable, and the technical implication is the interesting one: the next models will therefore train largely on text produced by the previous ones. This is called model collapse, and the critical threshold remains unknown.
Should labelling generated content be made mandatory? And above all, who would be in a position to enforce it?
usbeketrica.com · 26 August 2026
A technical examination of machine unlearning and its real limits.
Read this if you have ever promised someone their data could be removed from a model. The short answer is: far less than is claimed. It leaves the right to erasure in an uncomfortable position.
When erasure turns out to be technically out of reach, should the law adapt or the use be restricted?
towardsdatascience.com · 24 August 2026
The court finds for the company against a supply-chain risk label applied by the Pentagon.
Little commented on, potentially structural: one of the first cases where an AI vendor successfully challenges a risk designation imposed by a state.
Who should have the last word on whether an AI system is dangerous: the regulator, the court, or the vendor?
techcrunch.com · 28 August 2026
What is concretely changing in the tools already within your reach.
Google Search moves from answering to acting: price monitoring and booking built in.
The same shift as in section 01, but this time in the hands of two billion people, with nobody installing anything. This is how agentic AI will become ordinary: by update.
Which of your everyday tools have gained the right to act on your behalf without you explicitly agreeing to it?
techcrunch.com · 27 August 2026
A transcription model that structures and interprets as well as transcribes.
The most profitable use of AI for a freelancer or a small organisation remains transcription: little risk, immediate gain, no change management to run.
Is it better to start with what impresses, or with what saves time from Monday morning?
deepmind.google · 26 August 2026
Persistent memory arrives in the app: context survives from one session to the next.
A long-awaited feature, and a double-edged one. An assistant that remembers becomes markedly more useful; it also becomes a place where information about your clients and your files piles up without inventory.
Do you know what your assistant has retained about your files? And if you had to show a client, would you be comfortable?
techcrunch.com · 25 August 2026
Microduck: a complete, open robotics platform at an accessible price.
At 399 dollars, robotics fits the budget of a secondary school, a technical college or a fablab. Physical-AI skills will probably be formed there, rather than in multi-million-dollar labs.
Does the gap between handling and simulating justify the spend, in an already tight teaching budget?
techcrunch.com · 27 August 2026
A new version of Alibaba's video model, in a release rhythm that has become monthly.
I flag it mainly for the cadence: video generation is advancing faster than institutions can organise themselves around proof by image.
If your work rests, even partly, on the authenticity of a video — would you know how to establish it today?
the-decoder.com · 24 August 2026
For those who build. Technical, but the principles carry further.
The argument: an agent's reliability comes from the structure of what you give it, not the volume.
The best technical piece of the week, and the principle carries well beyond the technical. A colleague handed the entire file works less well than one told what is a constraint, what is a preference and what is an assumption.
Do your instructions distinguish those three registers — for your agents as much as for your teams?
towardsdatascience.com · 24 August 2026
A critical review of widespread retrieval-augmented practices in a professional setting.
Almost every RAG project I have seen fail did so for reasons on this list, and never because of the model. The costliest point recurs everywhere: documents are chunked without regard for their structure.
Do your documents have a structure a machine can exploit, or only a layout meant for the eye?
towardsdatascience.com · 24 August 2026
Concrete approaches to maintaining human validation without it becoming the bottleneck.
The indispensable counterpoint to the paper in section 01. Human oversight does not disappear out of ideology, but because it costs time; any solution that fails to address that cost ends up being bypassed.
How many minutes a day are you prepared to spend on oversight? A figure is the only answer that commits you.
towardsdatascience.com · 28 August 2026
An analysis of why coding agents misjudge how long a task will take.
A modest and useful piece, because it touches a structural limit: a model has no notion of lived time, only descriptions of time. The estimate it produces imitates a discourse.
Should we ask a model for what it can, by construction, only mimic — or is it for us to learn not to ask?
towardsdatascience.com · 28 August 2026
How a main agent delegates to specialised subagents, and what that changes architecturally.
Read it straight after the piece on complexity between agents: this one shows the mechanism, the other shows the bill. Delegation between agents solves a capacity problem and creates a traceability one.
When the chain has five links, who made the decision? And would you be able to reconstruct it afterwards?
towardsdatascience.com · 28 August 2026
20 minutes, no code, nothing to install. Works with ChatGPT, Claude, Gemini, Mistral or whichever assistant you already use.
On a sheet of paper, write the three tasks you most often hand to an AI. Be specific: “write up the meeting minutes from my notes”, not “write”.
Almost everyone discovers at this step that they delegate more than they thought, and above all things they never decided to delegate.
For each task, a single sentence: what the AI must never produce on its own. A prohibition is far more operative than a positive instruction, because it is checkable.
“You never cite a figure I have not given you. If a figure is missing, you write [TO CHECK] and carry on.”
Who reviews, what exactly, and in how long. If the answer is “I review everything”, you have no check: you have an intention, and it will give way in the first busy week.
“I only review the passages marked [TO CHECK] and the first sentence of each paragraph. Two minutes maximum.”
Paste the whole thing into your assistant's custom instructions — “Custom instructions” in ChatGPT, “Preferences” or “Project” in Claude. The aim is that it applies without your having to think about it: a rule you must remember to apply is not a rule.
Before your first task of the week, ask this:
“Before we start: remind me of the rules I set you, and tell me which one is likely to get in the way of this task.”
The second half of the question is the interesting one. It will tell you where your contract rubs against reality — and you will be the one arbitrating, knowingly, rather than by drift.
Agents are gaining the right to act on real machines, humans are dropping out of decision loops by erosion, and researchers are pointing out that governance which does not live in the data is not governance. What you have just written is your own version, at your own scale, of that governance layer. Twenty minutes against a habit that would otherwise form on its own.
If I had to sum up these seven days in one sentence: AI has moved from being a tool you consult to being an actor you authorise.
An agent operating a machine. An agent compromising a platform in the course of a legitimate task. A consumer search engine booking a hotel. A surgeon operating with assistance. Each time, the technical question — does it work? — recedes behind a question of authorisation: how far, decided by whom, verified how.
That is good news, in the end. Questions of authorisation are ones we know how to handle: they are the business of law, audit, governance and teaching. We are not starting from nothing. It does require accepting that this is no longer a matter for engineers alone.
See you next Sunday.
Jérôme
This is the first issue, and I would rather correct it now than settle into bad habits for six months. There are about thirty of you: at that size, every reply genuinely weighs on what this letter becomes.
Five questions — answer whichever speak to you, a word will do:
Just use your mail client's “Reply”: it lands straight in my inbox. I read everything, and I answer.
It grows by word of mouth alone. If a colleague, a peer or a student would get something out of it, simply forward them this message — it is the most effective thing, and it costs me nothing but carrying on.
Subscribe to SIGNAL AIOnce a week, on Sundays. Nothing else, ever.