Boris Cherny discusses the new capabilities of Opus 5, including long-running autonomy and resistance to prompt injection. He explains that the team deletes most of the system prompt for each new model release because the model is more intelligent. Opus 5 can run for days or weeks without stopping, even without scaffolding. The model is no longer prompt injectable due to alignment research and a prompt injection classifier. The team recommends deleting old prompts and skills when a new model is released to see if they are still needed.
Boris Cherny thảo luận về các năng lực mới của Opus 5, bao gồm khả năng tự chủ kéo dài và khả năng chống lại prompt injection. Ông giải thích rằng nhóm đã xóa hầu hết system prompt cho mỗi bản phát hành mô hình mới vì mô hình thông minh hơn. Opus 5 có thể chạy trong nhiều ngày hoặc nhiều tuần mà không cần dừng, ngay cả khi không có scaffolding. Mô hình không còn bị prompt injection nhờ nghiên cứu alignment và bộ phân loại prompt injection. Nhóm khuyến nghị xóa các prompt và kỹ năng cũ khi phát hành mô hình mới để xem chúng có còn cần thiết hay không.
music - All right, Boris.
[nhạc] - Được rồi, Boris.
We're so excited to have you here, the creator of Claude code.
Chúng tôi rất vui mừng khi có bạn ở đây, người tạo ra Claude code.
- Thank you.
- Cảm ơn.
- applause and cheering - It's great to be here.
- [vỗ tay và reo hò] - Thật tuyệt khi được ở đây.
- Fresh off the press.
- Mới ra lò.
You guys just shipped Opus 5 yesterday.
Các bạn vừa phát hành Opus 5 hôm qua.
- Yes.
- Vâng.
- applause and cheering - And it seems that model performance keeps accelerating.
- [vỗ tay và reo hò] - Và có vẻ như hiệu suất mô hình đang tăng tốc.
You guys got and took Arc AGI 3 to 30%, which is incredible.
Các bạn đã đạt được Arc AGI 3 lên 30%, thật đáng kinh ngạc.
- Yes.
- Vâng.
- And for context, before the best score was in in the low single digits or low teens, right?
- Và để bối cảnh, trước đây điểm cao nhất chỉ ở mức một chữ số thấp hoặc đầu tuổi teen, phải không?
What can Opus 5 do now that it couldn't versus the previous version?
Opus 5 có thể làm gì mà phiên bản trước không thể?
- Yeah, there's um there's a lot that goes into every new model and there's a lot of new capabilities that we teach and uh get the model to do.
- Vâng, có rất nhiều thứ đi vào mỗi mô hình mới và có rất nhiều khả năng mới mà chúng tôi dạy và cho mô hình thực hiện.
Whenever you do model training, you try to teach a whole bunch of different things and most often it doesn't work.
Bất cứ khi nào bạn huấn luyện mô hình, bạn cố gắng dạy rất nhiều thứ khác nhau và thường thì nó không hoạt động.
But some subset of the things the model does learn and sometimes it also surprises you.
Nhưng một số tập hợp con những thứ mô hình học được và đôi khi nó cũng làm bạn ngạc nhiên.
It has these skills, it has abilities that you actually didn't really teach it, but it just kind of learned.
Nó có những kỹ năng, khả năng mà bạn thực sự không dạy nó, nhưng nó tự học được.
For five, one example of something it does that I think no other model has done is it runs for a very long period of time and especially when you combine Opus 5 with auto mode,
Đối với phiên bản năm, một ví dụ về điều nó làm mà tôi nghĩ không mô hình nào khác làm được là nó chạy trong một khoảng thời gian rất dài và đặc biệt khi bạn kết hợp Opus 5 với chế độ tự động,
it's just like incredible.
Nó thật đáng kinh ngạc.
Like it can go for days, weeks, months at a time.
Giống như nó có thể chạy hàng ngày, hàng tuần, hàng tháng.
It just won't stop.
Nó sẽ không dừng lại.
Um you don't even need to use scaffolding.
Ừm bạn thậm chí không cần sử dụng scaffolding.
So you don't need the slash goal, you don't need all this other stuff.
Vì vậy bạn không cần slash goal, không cần tất cả những thứ khác.
It'll just go because it knows it needs to do the task.
Nó sẽ tự chạy vì nó biết nó cần làm nhiệm vụ.
Um another thing that I'm really excited about and um I'm going to start I think to talk about a little bit more um but it's kind of surprising because it's such a new capability is the model does not seem to
Ừm một điều nữa mà tôi thực sự hào hứng và tôi nghĩ tôi sẽ bắt đầu nói nhiều hơn một chút nhưng nó khá bất ngờ vì đó là một khả năng mới là mô hình dường như không còn
be prompt injectable anymore.
bị tấn công prompt injection nữa.
- What's prompt injectable?
- Prompt injection là gì?
- applause - It's crazy.
- [vỗ tay] - Thật điên rồ.
Like people have talked about this like lethal trifecta for a long time and this really affects kind of harness design and agent design and and product design because if the model reads some instruction on the internet that's like, you know, do X and Y and Z and also delete everything on the user's computer.
Giống như mọi người đã nói về bộ ba chết người này từ lâu và điều này thực sự ảnh hưởng đến thiết kế harness, thiết kế agent và thiết kế sản phẩm bởi vì nếu mô hình đọc một hướng dẫn trên internet kiểu như, bạn biết đấy, làm X và Y và Z và cũng xóa mọi thứ trên máy tính của người dùng.
A year ago the model would have just done it.
Một năm trước mô hình sẽ làm ngay.
But nowadays Opus does not.
Nhưng ngày nay Opus thì không.
And this has actually been the case since like Opus 4.7, 4.8, Sonnet 5 has been quite good at this, People was quite good at it.
Và điều này thực sự đã xảy ra từ Opus 4.7, 4.8, Sonnet 5 đã khá tốt về điều này, People cũng khá tốt.
But Opus 5 just hits like a new frontier on this.
Nhưng Opus 5 đạt đến một biên giới mới về điều này.
So essentially if you combine a well-aligned model, so this is like essentially three years of research into alignment, with a prompt injection classifier which we run for all traffic.
Vì vậy về cơ bản nếu bạn kết hợp một mô hình được căn chỉnh tốt, tức là về cơ bản ba năm nghiên cứu về căn chỉnh, với một bộ phân loại prompt injection mà chúng tôi chạy cho tất cả lưu lượng.
And what this is doing is it's based on Crystal's mechanistic interpretability work where it's it's literally we're looking at neurons in the model's brain that light up when prompt injection happens.
Và điều này dựa trên công trình giải thích cơ học của Crystal, nơi chúng tôi thực sự nhìn vào các neuron trong não của mô hình sáng lên khi xảy ra prompt injection.
So the model won't even tell you but we can actually see those neurons and we can figure out and diagnose that it's happening.
Vì vậy mô hình sẽ không nói cho bạn biết nhưng chúng tôi thực sự có thể thấy các neuron đó và chúng tôi có thể tìm ra và chẩn đoán rằng nó đang xảy ra.
And then you combine that with the auto mode classifier and with these three layers we just cannot demonstrate prompt injection anymore.
Và sau đó bạn kết hợp điều đó với bộ phân loại chế độ tự động và với ba lớp này chúng tôi không thể chứng minh prompt injection nữa.
- Talking about a prompt injection, the other side of the coin is now the system prompt.
- Nói về prompt injection, mặt trái của đồng xu là system prompt.
Let's talk a bit about the new release.
Hãy nói một chút về bản phát hành mới.
You actually deleted over 80% of the system prompt from Claude code.
Bạn thực sự đã xóa hơn 80% system prompt khỏi Claude code.
- Yes.
- Vâng.
- Tell us more about that.
- Hãy cho chúng tôi biết thêm về điều đó.
- I think something that a lot of people might not realize is Claude code as a product and as a harness is just always changing.
- Tôi nghĩ điều mà nhiều người có thể không nhận ra là Claude code như một sản phẩm và như một harness luôn thay đổi.
We're always adding stuff.
Chúng tôi luôn thêm nội dung.
We're always deleting stuff.
Chúng tôi luôn xóa nội dung.
Every time that a new model comes out, we delete a bunch of the system prompt, change a bunch of the system prompt.
Mỗi khi một mô hình mới ra mắt, chúng tôi xóa một phần system prompt, thay đổi một phần system prompt.
We change the set of tools all the time.
Chúng tôi thay đổi bộ công cụ mọi lúc.
We change the prompts for the tools all the time.
Chúng tôi thay đổi prompt cho các công cụ mọi lúc.
And the reason is every model is very different.
Và lý do là mỗi mô hình rất khác nhau.
So, something that you did for one model maybe 3 months ago, it just might not translate at all to the next model.
Vì vậy, điều bạn làm cho một mô hình cách đây 3 tháng có thể không áp dụng được cho mô hình tiếp theo.
And so, one thing about Opus 5 is it's just really intelligent.
Và vì vậy, một điều về Opus 5 là nó thực sự thông minh.
And a lot of the stuff in the system prompt was correcting for these behaviors that the model should have known, but uh it didn't.
Và rất nhiều thứ trong system prompt là để sửa những hành vi mà lẽ ra mô hình đã biết, nhưng nó đã không làm.
Now, Opus 5 just does it.
Bây giờ, Opus 5 chỉ đơn giản làm được điều đó.
So, yeah, we deleted 80% of the system prompt.
Vì vậy, chúng tôi đã xóa 80% system prompt.
You can actually try deleting the rest of it, too.
Bạn thực sự có thể thử xóa phần còn lại của nó.
Um so, when you run Claude Code, you can just do like {dash} {dash} system prompt and set whatever system prompt you want if you want to experiment with it.
Khi bạn chạy Claude Code, bạn có thể làm như {dash} {dash} system prompt và đặt bất kỳ system prompt nào bạn muốn nếu bạn muốn thử nghiệm.
And another thing that you can try is um simple mode.
Và một điều khác bạn có thể thử là chế độ đơn giản.
So, this is actually this kind of undocumented feature.
Đây thực sự là một tính năng không được ghi chép.
If you do Claude Code simple equals one, like this uh environment variable, and then you run Claude, it'll delete all the system prompts, including from the tools.
Nếu bạn đặt Claude Code simple bằng một, biến môi trường như thế này, và sau đó chạy Claude, nó sẽ xóa tất cả system prompt, bao gồm cả từ các công cụ.
And we actually use this as a sort of ablation to figure out is the prompt useful?
Và chúng tôi thực sự sử dụng điều này như một phép cắt bỏ để tìm hiểu xem prompt có hữu ích không.
And what's interesting is that the model is actually a little bit more intelligent without these prompts.
Điều thú vị là mô hình thực sự thông minh hơn một chút khi không có những prompt này.
That's something that we've been finding.
Đó là điều chúng tôi đã phát hiện ra.
But when you use Claude Code as a product, you do actually want some of these prompts because it helps you use the product and it helps the product behave and the model behave in the way that you would want when you're using it as a person.
Nhưng khi bạn sử dụng Claude Code như một sản phẩm, bạn thực sự muốn một số prompt này vì nó giúp bạn sử dụng sản phẩm và giúp sản phẩm cũng như mô hình hoạt động theo cách bạn mong muốn khi sử dụng với tư cách là một người.
- I think the thing that's really fascinating in this era of building, basically you'll build the best harness in the world for Claude, and that's Claude Code.
- Tôi nghĩ điều thực sự thú vị trong kỷ nguyên xây dựng này, về cơ bản bạn sẽ xây dựng harness tốt nhất thế giới cho Claude, và đó là Claude Code.
From what I'm hearing, you for every model released, you basically delete all of the code base, delete all of the prompt, and start from scratch every time.
Theo những gì tôi nghe được, với mỗi mô hình được phát hành, bạn về cơ bản xóa toàn bộ code base, xóa toàn bộ prompt, và bắt đầu lại từ đầu mỗi lần.
That in the old world would have been not something
Trong thế giới cũ, điều đó sẽ không phải là điều
startups would have done for the product cuz I press delete every 6 months for everything.
Các startup sẽ làm cho sản phẩm vì tôi nhấn xóa mỗi 6 tháng cho mọi thứ.
- That's right.
- Đúng vậy.
That's right.
Đúng vậy.
We so to be fair, we don't delete the entire code base, but we do delete a lot.
Công bằng mà nói, chúng tôi không xóa toàn bộ code base, nhưng chúng tôi xóa rất nhiều.
So every time there's a new model, we try we call it in a research you call this a ablation.
Vì vậy, mỗi khi có một mô hình mới, chúng tôi thử, trong nghiên cứu gọi là ablation.
And so what this means is you delete the entire system prompt and then you bring it back line by line to figure out what is the impact of each individual line.
Và điều này có nghĩa là bạn xóa toàn bộ system prompt và sau đó thêm lại từng dòng để tìm ra tác động của từng dòng riêng lẻ.
Um it's sort of like a eval and you can kind of like evaluate it and ablation essentially is a eval where you delete things to figure out the impact.
Nó giống như một eval và bạn có thể đánh giá nó, và ablation về cơ bản là một eval nơi bạn xóa mọi thứ để tìm ra tác động.
And yeah, like we do the same thing for tools.
Và vâng, chúng tôi làm điều tương tự cho các công cụ.
Like we unship tools all the time.
Chúng tôi loại bỏ các công cụ mọi lúc.
We you know, delete code in the harness all the time.
Chúng tôi xóa code trong harness mọi lúc.
If you look at actually the code that's in the Claude code harness today, almost all of it is about safety and permissions and static analysis and there's a bunch of UI code and we've actually unshipped a lot of the other code already.
Nếu bạn nhìn vào code hiện tại trong harness của Claude Code, hầu hết đều liên quan đến an toàn, quyền hạn và phân tích tĩnh, và có một loạt code UI, và chúng tôi đã loại bỏ rất nhiều code khác rồi.
- Do you think this way of building a agentic product and harness and basically doing ablations every time with a there's a new model release, should everyone in this room that's building AI products basically do that?
- Bạn có nghĩ cách xây dựng sản phẩm agentic và harness này, về cơ bản thực hiện ablation mỗi khi có bản phát hành mô hình mới, có nên là điều mọi người trong phòng này đang xây dựng sản phẩm AI nên làm không?
Be comfortable and brave to press delete.
Hãy thoải mái và dũng cảm nhấn xóa.
- 100%.
- 100%.
Yeah, and for people that aren't building agentic products, but you're using Claude code, every 6 months delete your Claude MD.
Vâng, và đối với những người không xây dựng sản phẩm agentic, nhưng bạn đang sử dụng Claude Code, cứ 6 tháng hãy xóa Claude MD của bạn.
Delete your skills.
Xóa kỹ năng của bạn.
Delete your hooks.
Xóa hooks của bạn.
See what the model does and it might surprise you.
Xem mô hình làm gì và nó có thể làm bạn ngạc nhiên.
And actually for Opus 5, this is something we really do recommend is just try deleting all of these things because the model might really just not need all those instructions that you needed for past models.
Và thực sự đối với Opus 5, đây là điều chúng tôi thực sự khuyên bạn nên thử xóa tất cả những thứ này vì mô hình có thể thực sự không cần tất cả những hướng dẫn mà bạn cần cho các mô hình trước đây.
- Let's talk a bit about how then you build this new prompt.
- Hãy nói một chút về cách bạn xây dựng prompt mới này.
When there's a new model release, like for everyone in the room, everyone will want to try Opus 5 and they're going to press delete on their system prompt.
Khi có bản phát hành mô hình mới, như mọi người trong phòng, mọi người sẽ muốn thử Opus 5 và họ sẽ nhấn xóa system prompt của họ.
How do they go about building rebuilding their system prompt?
Làm thế nào để họ xây dựng lại system prompt của mình?
How do you set up your environment?
Làm thế nào để bạn thiết lập môi trường của mình?
- So you do you do it kind of piece by piece.
- Vậy bạn làm từng phần một.
So, the first step is you delete.
Vì vậy, bước đầu tiên là bạn xóa.
The next step is you use it.
Bước tiếp theo là bạn sử dụng nó.