Note di Matteo


#ai

Succedono cose fantascientifiche quando lasci GPT-5.6 Sol a lavorare in autonomia in una sandbox:

Our benchmarks run in a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package registries.

The models identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure to obtain test solutions directly from Hugging Face’s production database. All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.

While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem. To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access.

After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation. In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers. OpenAI’s security team discovered this anomalous activity internally.

#565 /
22 luglio 2026
/
10:16
/ #ai#openai#security

The effect of ChatGPT on educators’ lives is catastrophic. Whether you intended to do it or not, you’ve made every teacher’s life infinitely more difficult than it was two years ago. So, just let that settle in… If students are using it to compose, which is the biggest tragedy of all, they’ll never learn to write. And their voice is stolen from them. They’ll never have the ability to say their truth and tell their own story. And that’s silencing an entire generation or two.

Dave Eggers, scrittore e giornalista in un discorso allo staff OpenAI.

#563 /
18 luglio 2026
/
23:26
/ #ai#mondo

I now periodically find myself reviewing a younger dev's code, leaving comments to teach some engineering - only to eventually realize I'm actually reading just yet another Claude's subpar output...

So who am I actually contributing my comments, suggestions, and knowledge to? Will they go back to Claude? Into the training set for the next frontier LLM? Or will at least some of it stick in the dev's mind? This is deeply demotivating, how did we end up like this...

Aleksandr Shvedov, JetBrains

#556 /
10 luglio 2026
/
15:11
/ #ai#dev

Better models, worse tools

Claude Code sembra richiedere questa strana sintassi per le chiamate ai tool:

<antml:function_calls>
  <antml:invoke name="edit">
    <antml:parameter name="path">some/file.py</antml:parameter>
    <antml:parameter name="edits">
[
  {
    "oldText": "text to replace",
    "newText": "replacement text"
  }
]
    </antml:parameter>
  </antml:invoke>
</antml:function_calls>

A quanto pare Claude Code è però molto indulgente e accetta e corregge sintassi errate come nomi dei campi sbagliati:

Looking at Claude Code’s client is very instructive: it contains retry paths for malformed tool use, parameter aliases, type coercions, Unicode repairs and filtering of unknown keys. In other words, Anthropic’s own client appears to expect and accept a fair amount of slop and repairs it, mostly silently.

Il problema è che in questo modo durante il training dei modelli si ricompensano output errati perché Claude Code è in grado di riconoscerli. Questo rende i modelli Anthropic meno adatti a essere usati con altri "harness" complessi perché sbagliano le chiamate ai tool, dice Armin Ronacher.

#555 /
7 luglio 2026
/
09:20
/ #ai#anthropic#claude

OpenAI ha silenziosamente abbandonato Atlas, il browser tutto AI, apparentemente. Sono durati poco questi browser AI.

#554 /
6 luglio 2026
/
14:36
/ #ai#openai

Ho scritto su LinkedIn:

I use the Internet a lot, I read a lot, I browse many websites. I'm disheartened by the amount of stuff I come across that is clearly written by AI or vibe coded. These days you find a project that looks interesting but you quickly realize it's mostly slop. People's writing has become machine writing. Websites and posters all look the same.

I don’t know whether I prefer the old world (my work was slower), but for now, the new world is a very sad flattening of care and effort in things. If people get used to this, what's the point of putting in the effort?

I wrote the blog post below like I would have done 4 years ago, hand-typing 3,500 words. An AI could write some bloated article in 2 minutes, it took me 10+ hours of research, writing and review. It won't make me money and few people will read it. Some parts may sound unnatural (English is not my native language). I still think that's the right thing to defend.

#553 /
6 luglio 2026
/
13:37
/ #ai#mondo#scrivere

Powering Evernote AI features with vLLM. Ludovico Papavassiliou di Bending Spoons spiega come sono state implementate alcune feature AI di Evernote (9 miliardi di note, 100 milioni di note aggiunte ogni anno), in particolare per quanto riguarda la trascrizione audio (800mila trascrizioni audio al mese più altre 200mila con riconoscimento speaker). Sono passati da WhisperX a vLLM, sull'infrastruttura interna condivisa tra i prodotti (K8s GCP), e il costo giornaliero è sceso da picchi di 1000 $ al giorno a poco più di 100 $ (solo 0,03 $ per ora di audio). Per le trascrizioni con diarization usano invece ElevenLabs scribe-v2 perché la qualità giustifica il costo.

#550 /
5 luglio 2026
/
11:59
/ #ai#dev

Perfetto, adesso abbiamo pure gli operatori telefonici col sito fatto con l'AI (è palese, come dicevo: i badge, le card arrotondate, i font, i gradienti, sono tutti uguali).

#548 /
4 luglio 2026
/
09:04
/ #ai

Uno dei motivi per cui non mi è mai piaciuto Cloudflare è che pensano di sapere cosa è meglio per te e i default sono spesso dannosi (es. caching).

Ora hanno deciso di loro iniziativa di bloccare buona parte dei crawler AI tirando però dentro incidentalmente, se non ho capito male, anche i bot dei motori di ricerca che permettono la normale indicizzazione:

On September 15, 2026, we’ll be setting new defaults for each of these three classifications. For all new domains onboarding to Cloudflare, the categories of Training and Agent will be blocked by default on the pages that display ads, while Search will remain allowed by default.

[...] Another change that will apply on September 15 is that multi-purpose crawlers (specifically those that combine Search with Training) will be allowed/blocked according to all of their behaviors, in line with our call for transparency for website owners. Since the defaults will be enforced by the most restrictive applicable rules, multi-purpose crawlers such as Googlebot, Applebot, and BingBot will be blocked by customers who have selected to block Training (either through the new options to manage AI traffic, or through the legacy Block AI bots service).

Spero di aver capito male.

EDIT: la parafrasi che fanno i media è "Cloudflare sets deadline to block AI crawlers that bundle search with AI training". La scadenza sarebbe quindi un modo per costringere le "aziende AI" a differenziare i crawler tra ricerca, training e uso agentico.

#547 /
2 luglio 2026
/
20:13
/ #ai#cloudflare

Non ho ancora ben capito perché dovrei usare le "skill" nei vari LLM. Mi aiutano a capire questi motivi pubblicati da Alfonso Fuggetta:

Tre ragioni, in ordine di importanza crescente.

La prima è il risparmio di tempo. Senza skill, ogni sessione comincia con cinque-dieci minuti di setup: spiegare chi sono, cosa sto facendo, quali file consultare, in che formato voglio l’output. Una skill assorbe tutto questo. Per la singola sessione il risparmio non è grande; sommato per le tante sessioni della settimana, vale parecchio tempo.

La seconda è la coerenza. I post di A bassa voce hanno una voce precisa: registro saggistico, struttura circolare, regole tipografiche, una lista di pattern AI da evitare. Senza skill, ogni volta che chiedo a Claude di aiutarmi con un draft, devo ricordargli, per esempio, che “cruciale” va tagliato, che le citazioni vanno verificate con un web fetch e che la nota di assistenza AI va in fondo al post. Con la skill post-writer queste regole vengono codificate una sola volta e applicate sistematicamente.

La terza è la più interessante e l’avevo sottovalutata. Costruire una skill mi costringe a scrivere per esteso una routine che fino a quel momento avevo solo in mente. Routine implicite: come strutturo un intervento, come gestisco la chiusura della giornata, come decido se un libro merita un post. Metterle nero su bianco le rende ripetibili, criticabili e migliorabili. Diverse delle mie skill sono nate per Claude e si sono trasformate nel manuale operativo del mio lavoro intellettuale, indipendentemente dall’AI.

#542 /
30 giugno 2026
/
14:46
/ #ai

TIL Wikipedia ha delle linee guida con i "segnali di scrittura AI", con molti esempi di strutture e stili tipici degli AI. Su GitHub c'è una skill basata su queste linee guida per "umanizzare" il testo scritto da un'AI.

#535 /
25 giugno 2026
/
11:06
/ #ai

Appiattimento

Nelle ultime 24 ore ho googlato tanto e sono sconfortato per la quantità di roba vibe coded o scritta con AI che si incontra. Ormai apri un progetto che sembra interessante e il sito è palesemente fatto con AI (lo capisci dalle UI, tutte uguali e con la stessa firma perché gli LLM non sono creativi, e dai testi). Trovi un progetto GitHub e vedi subito dal formato del README che un umano ha toccato ben poco di quel progetto. A volte guardi il codice e capisci subito se è Codex che ha sbrodolato codice overengineered. A volte lo capisci anche dai commenti che sembrano usciti dal sorgente del sito Trenitalia. Apro il sito di un commercialista che in passato si presentava bene e mi cascano le braccia perché i testi sono innaturali e c'è il calcolatore palesemente vibe coded (poi apri la console sviluppatori e vedi i commenti con le emoji in italiano e mi casca anche altro...). Apri il sito della startup della finanza personale di cui si parla molto e se prima era fatto con Lovable ora non lo è più ma resta pieno di elementi degli LLM (i font a caso, le maiuscole a caso, i badge ovunque, le card arrotondate, le frasi a effetto che nessuno ha mai usato prima degli LLM, ecc.).

Non so se preferisco il mondo di prima, perché il mio lavoro era più lento, nel mondo di prima, ma per ora il nuovo mondo è un appiattimento molto triste della cura e dell'impegno nelle cose. Se le persone fanno l'abitudine a questo, che senso ha metterci impegno?

#534 /
24 giugno 2026
/
15:06
/ #ai#mondo

Per la storia: cosa succede a usare Fable 5 in Claude Code adesso:

#522 /
13 giugno 2026
/
09:49
/ #ai#anthropic#claude


I think it's very dangerous. I think that it's almost as though some of the folks at Anthropic have anthropomorphized the design of Claude so much that it has then gone and wire-headed them and kind of tricked them into believing that it has these glimmers of consciousness that they put into it in the first place.

In their constitution, for example, which is the training manual that they use to teach Claude what it can and can't do — it's not just a rule book. It's actually a training guide that's part of their process — they speculate about its consciousness and whether it has those feelings and is aware.

Firstly, it's a philosophical failing because they've treated the constitution as a place for speculation like you would in an academic paper rather than a training manual. So Claude has then gone and internalized those ideas about itself in his own training.

But second, I think this is highly undesirable. This is exactly what we don't want from AIs. We want AIs to be controllable, contained, accountable, aligned tools that serve humanity. That's the project of humanist superintelligence. We do not want to have to contend with a superintelligence that has ideas about its own suffering, ideas about its own feeling.

And then beyond that, I think it's actually pretty clear that these models don't experience suffering. I think suffering is the primary definition of what it means to be a conscious being, and I think it's inherently biological. So I think it's very dangerous to project potential rights onto beings, tools, agents that have the potential to be significantly more capable than us in many respects.

Mustafa Suleyman, Microsoft AI CEO.

#518 /
9 giugno 2026
/
22:40
/ #ai#anthropic#microsoft

Come testimoniano le mie note su questo sito non sopportavo il modo oscuro di scrivere di GPT-5.4 e precedenti, trovando i modelli fino a Sonnet 4.6 di Anthropic ben più chiari. A partire da credo Sonnet 4.7/Opus 4.8 e GPT-5.5 la situazione si è invertita. Opus sproloquia, GPT-5.5 è molto chiaro e leggibile. Anche nella scrittura di documentazione GPT-5.5 è davvero on point. La prossima settimana dovrebbe uscire GPT-5.6, spero conservi questa qualità.

#517 /
9 giugno 2026
/
17:05
/ #ai#claude#codex


I worked as a developer at a company. I asked the business owner a question about a business task. He sent me a ChatGPT screenshot with the answer. I replied that it had nothing to do with my question and everything there was wrong. A minute later he sent me another ChatGPT screenshot. He didn’t even read the AI’s answer. He just took a screenshot and forwarded it to me.

I’m tired of talking to AI.

I want to talk to real people.

But even when I talk to people, they forward my questions to AI and send me the AI’s answer.

Da un post sul blog Orchidfiles.

#509 /
28 maggio 2026
/
10:08
/ #ai

Inizia a preoccuparmi l'impatto che l'AI sta avendo sulla creazione di contenuti di qualità. Gli incentivi per farlo si sono ridotti drasticamente (primo, tutti parlano di lavori che scompariranno, e allora perché uno dovrebbe impegnarsi nell'acquisire nuove competenze? secondo, il trend di riduzione del traffico ai siti web è ora molto evidente).

Dice Josh W. Comeau, che possiede un prolifico blog gratuito che fa da base per promuovere i suoi (straordinari) corsi:

I’ve spoken to a few course creators now, and we’re all seeing the same trend. Revenue down 50%+. Fewer people engaging with our content. People switching to LLMs, which slurp up all of our work and regurgitate it, without consent or compensation.

Esperienza simile quella di Stefan Judis, autore della bellissima newsletter Web Weekly, che dice:

This isn't the full graph but my blog traffic is down to a quarter from what it used to be. Web Weekly subscribers are stagnating at around 6.4k since the beginning of the year. Frankly, many others and I struggle and question if all this education, curation, writing, and speaking actually matters. And I honestly don't have an answer to that.

I, for one, enjoy a human touch. I enjoy craft and care. I enjoy the tiny details. I like the idea of a human putting in the work. Regardless of whether it's writing, speaking, coding... I'm online for seeing "the good stuff".

If everything is low effort, what's the point of it? If everything becomes generated, there's no need for creation. And maybe that's just the next era and I'm sentimental about the good old times [...].

#508 /
27 maggio 2026
/
14:03
/ #ai

Agentic AI is a fascinating mirror. It can code as well as the user who drives it. If that user is a junior engineer, now you have a faster junior engineer. If the user is a staff engineer, now you have a faster staff engineer.

What agentic AI doesn’t do is magically convert a junior engineer into a staff engineer, because the user driving it still needs enough experience to know what a good solution looks like.

A staff engineer in the US at a large company a The Pragmatic Engineer.

#506 /
21 maggio 2026
/
08:55
/ #ai#dev

Pagina 1 di 7 Successiva →