The video introduces StockPilot, an inventory management agent that has become overcomplicated with a 400-line system prompt, 12 tools, and sub-agents, leading to performance degradation. Evals show failures due to inefficient paths and communication breakdowns between sub-agents and the orchestrator. The workshop aims to improve the agent's architecture using proper agentic primitives like tools, skills, and sub-agents to restore performance. The current eval success rate is about 83%, which is unacceptable in manufacturing. The agent hallucinated a promotion multiplier due to context problems. The workshop will involve running evals, triaging issues, updating agent design, and hill climbing towards eval improvement.
Video giới thiệu StockPilot, một agent quản lý hàng tồn kho đã trở nên quá phức tạp với prompt hệ thống 400 dòng, 12 công cụ và các sub-agent, dẫn đến suy giảm hiệu suất. Các bài đánh giá cho thấy thất bại do đường dẫn không hiệu quả và sự cố giao tiếp giữa sub-agent và bộ điều phối. Hội thảo nhằm cải thiện kiến trúc agent bằng cách sử dụng các nguyên thủy agentic phù hợp như công cụ, kỹ năng và sub-agent để khôi phục hiệu suất. Tỷ lệ thành công hiện tại của eval là khoảng 83%, không thể chấp nhận được trong sản xuất. Agent đã ảo tưởng ra một hệ số khuyến mãi do vấn đề ngữ cảnh. Hội thảo sẽ bao gồm chạy eval, phân loại vấn đề, cập nhật thiết kế agent và leo đồi hướng tới cải thiện eval.
All right.
Được rồi.
Fantastic.
Tuyệt vời.
Can everyone hear me?
Mọi người nghe tôi nói rõ không?
Thumbs up.
Giơ ngón cái lên.
All good.
Tốt hết rồi.
All right, everyone.
Được rồi, mọi người.
I hope that you have had a fantastic day at Code with Claude London.
Tôi hy vọng các bạn đã có một ngày tuyệt vời tại Code with Claude London.
So far today, my name is Will.
Hôm nay, tôi là Will.
I'm on our engineering team at Anthropic.
Tôi thuộc đội ngũ kỹ thuật tại Anthropic.
I sit on a team called Applied AI.
Tôi làm việc trong nhóm có tên Applied AI.
What that means is I essentially split my time between internal engineering work and time spent building agents with customers.
Điều đó có nghĩa là tôi chủ yếu chia thời gian giữa công việc kỹ thuật nội bộ và xây dựng các agent với khách hàng.
So folks, imagine that you built and shipped an agent to solve a problem.
Vậy, hãy tưởng tượng bạn đã xây dựng và triển khai một agent để giải quyết một vấn đề.
I'm sure that's something that a lot of the folks in this room have actually done.
Tôi chắc rằng đó là điều mà nhiều người trong phòng này đã từng làm.
And imagine this that this agent worked fantastic, right?
Và hãy tưởng tượng rằng agent này hoạt động tuyệt vời, đúng không?
But it worked so well that a few weeks after shipping, you were asked to add some additional capability to the agent.
Nhưng nó hoạt động tốt đến nỗi vài tuần sau khi triển khai, bạn được yêu cầu thêm một số khả năng bổ sung cho agent.
A few weeks after that, you received more business requirements and you added additional capability.
Vài tuần sau đó, bạn nhận được thêm yêu cầu kinh doanh và bạn đã thêm khả năng bổ sung.
This pattern continued and continued until before you know it, your system prompt had grown to become
Mô hình này tiếp diễn và tiếp diễn cho đến khi bạn nhận ra, system prompt của bạn đã phát triển thành
several hundred lines long.
dài vài trăm dòng.
You have dozens of tools and sub agents that exist for your agent and because of the complexity you've started to see regressions in the areas that your agent was previously accelerating in.
Bạn có hàng tá công cụ và agent con cho agent của mình và vì sự phức tạp, bạn bắt đầu thấy sự suy giảm trong các lĩnh vực mà agent trước đây đã tăng tốc.
So if this is you, you're not alone.
Vậy nếu bạn rơi vào trường hợp này, bạn không đơn độc.
We see this type of scenario happen pretty commonly with customers and actually with ourselves included in that.
Chúng tôi thấy loại kịch bản này xảy ra khá phổ biến với khách hàng và thực tế là cả chúng tôi cũng vậy.
So within this
Vì vậy, trong
workshop, we are going to simulate an agent that has essentially grown to a complexity where we start to see degradation in its performance.
buổi workshop này, chúng ta sẽ mô phỏng một agent đã phát triển đến mức phức tạp khiến chúng ta bắt đầu thấy sự suy giảm hiệu suất.
We're then going to walk through some of the decisions that we as engineers and architects make in order to improve the design of our agent to restore the performance that we expect with the additional capability.
Sau đó, chúng ta sẽ xem xét một số quyết định mà chúng ta với tư cách là kỹ sư và kiến trúc sư đưa ra để cải thiện thiết kế agent nhằm khôi phục hiệu suất mong đợi với khả năng bổ sung.
Specifically,
Cụ thể,
we're going to make some decisions around tools and skills and sub agents.
Chúng ta sẽ đưa ra một số quyết định về công cụ, kỹ năng và agent con.
As we modernize the stack of our agent, we want to make sure that we're using the right agentic primitives at the right time.
Khi hiện đại hóa stack của agent, chúng ta muốn đảm bảo sử dụng đúng nguyên thủy agentic vào đúng thời điểm.
So, when do you use a tool?
Vậy, khi nào bạn sử dụng công cụ?
When do you use a skill?
Khi nào bạn sử dụng kỹ năng?
And when do you use a sub agent?
Và khi nào bạn sử dụng agent con?
We're going to talk through all of that in this session.
Chúng ta sẽ thảo luận tất cả những điều đó trong buổi này.
As I mentioned folks, this session will be hands-on.
Như tôi đã đề cập, buổi này sẽ mang tính thực hành.
So, let's go ahead and get started.
Vậy, chúng ta hãy bắt đầu.
I first want to walk you through our problem statement in our agent.
Trước tiên, tôi muốn giới thiệu cho các bạn về bài toán và agent của chúng ta.
So for the purposes of this session, we're going to be focusing on an agent called Stock Pilot.
Trong buổi này, chúng ta sẽ tập trung vào một agent có tên Stock Pilot.
This is an
Đây là một
inventory management agent that was designed by and for a midsize retailer.
agent quản lý hàng tồn kho được thiết kế bởi và cho một nhà bán lẻ cỡ vừa.
The agent that you see on the screen can do several things.
Agent mà bạn thấy trên màn hình có thể làm nhiều việc.
It can flag low levels of stock.
Nó có thể cảnh báo mức tồn kho thấp.
It can forecast demand.
Nó có thể dự báo nhu cầu.
It can pick suppliers.
Nó có thể chọn nhà cung cấp.
It can file POS and ultimately it can write weekly reports for the employees of this retailer.
Nó có thể ghi nhận POS và cuối cùng có thể viết báo cáo hàng tuần cho nhân viên của nhà bán lẻ này.
Now, none of these capabilities are particularly complex on their own, but again, the issue is that we've essentially bolted capabilities onto our agent over time without modernizing our architecture.
Không có khả năng nào trong số này đặc biệt phức tạp, nhưng vấn đề là chúng ta đã ghép thêm các khả năng vào agent theo thời gian mà không hiện đại hóa kiến trúc.
This complexity has started to cause some problems.
Sự phức tạp này đã bắt đầu gây ra một số vấn đề.
Let's take a look at the
Hãy xem
actual architecture today of the agent.
kiến trúc thực tế hiện tại của agent.
Folks, today the agent is facilitated by a single orchestrator.
Hiện tại, agent được điều phối bởi một bộ điều phối duy nhất.
So you see the stockpilot orchestrator sitting at the top of the screen.
Bạn thấy bộ điều phối stockpilot ở đầu màn hình.
The agent has a system prompt as I mentioned that's grown to be about 400 lines long.
Agent có một system prompt như tôi đã đề cập đã phát triển dài khoảng 400 dòng.
It has 12 different tools.
Nó có 12 công cụ khác nhau.
Three of those tools happen to be wrappers around sub agents with completely isolated context windows.
Ba trong số những công cụ đó là các wrapper xung quanh các tác tử phụ với các ngữ cảnh hoàn toàn biệt lập.
So if you have the repo pulled
Vì vậy, nếu bạn đã kéo kho lưu trữ về
up, which we'll go into more detail in just a bit, there's an agent that's under a folder called before, which essentially walks through this agent.
lên, chúng ta sẽ đi vào chi tiết hơn một chút, có một tác tử nằm trong thư mục có tên before, về cơ bản sẽ đi qua tác tử này.
Exactly.
Chính xác.
So again, orchestrator, long system prompt, a lot of tools, we have a lot of sub agents.
Một lần nữa, người điều phối, system prompt dài, nhiều công cụ, chúng ta có nhiều tác tử phụ.
The result of this is that our eval started to dip.
Kết quả là điểm đánh giá của chúng tôi bắt đầu giảm.
So let's imagine how we got here for a moment.
Vậy hãy tưởng tượng chúng ta đã đến đây như thế nào.
Again, like we built that agent
Một lần nữa, chúng tôi đã xây dựng tác tử đó
up front to solve a really specific problem.
ngay từ đầu để giải quyết một vấn đề rất cụ thể.
We received business requirements to say add maybe some forecasting capability to our inventory management agent.
Chúng tôi nhận được yêu cầu kinh doanh là thêm khả năng dự báo vào tác tử quản lý hàng tồn kho của mình.
So what we decided to do was essentially just spin up a forecaster as a sub agent.
Vì vậy, những gì chúng tôi quyết định làm là về cơ bản tạo ra một tác tử dự báo như một tác tử phụ.
Again later on we received more requirements to add report writing capability to our agent.
Sau đó, chúng tôi lại nhận thêm yêu cầu thêm khả năng viết báo cáo vào tác tử của mình.
So we decided to add another sub
Vì vậy, chúng tôi quyết định thêm một tác tử phụ
agent for that report writing capability.
khác cho khả năng viết báo cáo đó.
Again, our eval started to dip over time because we added more and more complexity while just bolting this capability on.
Một lần nữa, điểm đánh giá của chúng tôi bắt đầu giảm theo thời gian vì chúng tôi thêm ngày càng nhiều độ phức tạp trong khi chỉ ghép thêm khả năng này.
So, let's take just a little bit of time and talk about eval specifically for this agent.
Vì vậy, hãy dành một chút thời gian để nói về đánh giá cụ thể cho tác tử này.
Folks, we have 12 different eval tasks across five different types of graders.
Thưa các bạn, chúng tôi có 12 nhiệm vụ đánh giá khác nhau trên năm loại chấm điểm khác nhau.
So, my colleague gave a talk on eval shortly before this.
Đồng nghiệp của tôi đã có một bài nói về đánh giá ngay trước buổi này.
Evals will have a component
Đánh giá sẽ có một phần
within this workshop, but it won't be the main focus.
trong hội thảo này, nhưng nó sẽ không phải là trọng tâm chính.
I'll give you a quick summary of the tactical eval that we're using for this agent.
Tôi sẽ tóm tắt nhanh về đánh giá chiến thuật mà chúng tôi đang sử dụng cho tác tử này.
On the left side of the screen, you see some IDs.
Ở phía bên trái màn hình, bạn thấy một số ID.
You see several evals that start with the letter R.
Bạn thấy một số đánh giá bắt đầu bằng chữ R.
This stands for regression.
Điều này có nghĩa là hồi quy (regression).
These are more realistic single turn tasks that we grade the model's capability on.
Đây là những nhiệm vụ đơn lượt thực tế hơn mà chúng tôi chấm điểm khả năng của mô hình.
So imagine I give the model a task.
Hãy tưởng tượng tôi giao cho mô hình một nhiệm vụ.
The model comprehends that task in the for within the agent calls some tools and then provides a response back to me.
Mô hình hiểu nhiệm vụ đó, trong tác tử, gọi một số công cụ và sau đó cung cấp phản hồi lại cho tôi.
We're essentially
Về cơ bản chúng tôi
evaluating that response.
đánh giá phản hồi đó.
We also have some more complex tasks that we're grading the model on.
Chúng tôi cũng có một số nhiệm vụ phức tạp hơn mà chúng tôi đang chấm điểm mô hình.
So you see those F IDs, the IDs that start with F on the left side of the screen, that stands for failure mode.
Vì vậy, bạn thấy các ID F, các ID bắt đầu bằng F ở bên trái màn hình, viết tắt của failure mode (chế độ lỗi).
In this case, we're evaluating the model over a more complicated multi-turn task that we're grading.
Trong trường hợp này, chúng tôi đánh giá mô hình trên một nhiệm vụ đa lượt phức tạp hơn mà chúng tôi đang chấm điểm.
Now again I won't go into eval too specifically.
Bây giờ, tôi sẽ không đi quá chi tiết về đánh giá.
We have a number of
Chúng tôi có một số
different types of graders that are both deterministic and non-deterministic.
loại chấm điểm khác nhau, cả xác định và không xác định.
Right?
Đúng không?
When I talk about deterministic eval count and like latency and like the number of tokens that are used as our agent is completing a particular task and we're tracking those deterministic metrics over time.
Khi tôi nói về đánh giá xác định như số lượng token được sử dụng khi tác tử của chúng tôi hoàn thành một nhiệm vụ cụ thể và chúng tôi theo dõi các chỉ số xác định đó theo thời gian.
We're also using the idea of LLM as a judge to evaluate the non-deterministic characteristics of our agent.
Chúng tôi cũng sử dụng ý tưởng LLM làm người đánh giá để đánh giá các đặc tính không xác định của tác tử.
So, personality and tone and style and output quality.
Vì vậy, tính cách, giọng điệu, phong cách và chất lượng đầu ra.
We're using a nondeterministic grader as a part of our eval to evaluate our agents' non-deterministic characteristics.
Chúng tôi sử dụng một bộ chấm điểm không xác định như một phần của đánh giá để đánh giá các đặc tính không xác định của tác tử.
Now, we're going to run the evals for our agent in just a bit, but when you do, you'll find that the agent is struggling a bit.
Bây giờ, chúng tôi sẽ chạy đánh giá cho tác tử của mình một chút, nhưng khi bạn làm, bạn sẽ thấy rằng tác tử đang gặp khó khăn.
I'll talk about some of these eval in just a little bit more depth.
Tôi sẽ nói sâu hơn một chút về một số đánh giá này.
So F1 on the screen, third from the bottom.
Vậy F1 trên màn hình, thứ ba từ dưới lên.
This is essentially simulating a daily low stock sweep.
Điều này về cơ bản mô phỏng một đợt quét hàng tồn thấp hàng ngày.
So again, this is an inventory agent.
Một lần nữa, đây là một tác tử quản lý hàng tồn kho.
We're simulating our ability to look through
Chúng tôi mô phỏng khả năng xem xét
all of our inventory and pull the low levels of stock.
tất cả hàng tồn kho của chúng tôi và lấy các mức tồn kho thấp.
This eval will actually fail because the agent is going to do the right thing, but it's going to take a very winding path to do so.
Đánh giá này thực sự sẽ thất bại vì tác tử sẽ làm đúng, nhưng nó sẽ đi một con đường rất quanh co để làm điều đó.
So instead of taking the straightest line from point A to point B, the agent is going to take a very inefficient path.
Vì vậy, thay vì đi đường thẳng nhất từ điểm A đến điểm B, tác tử sẽ đi một con đường rất kém hiệu quả.
It's going to get to the right end, but it's going to fail the eval because it's not at the efficiency that we'd like.
Nó sẽ đến đích đúng, nhưng nó sẽ thất bại trong đánh giá vì nó không đạt được hiệu quả như chúng tôi mong muốn.
F2 on the
F2 trên
screen is another eval that you'll see fail.
màn hình là một bài đánh giá khác mà bạn sẽ thấy thất bại.
This evaluates the ordering process under a particular promotion package.
Bài đánh giá này đánh giá quy trình đặt hàng theo một gói khuyến mãi cụ thể.
This is going to fail because we are using a sub agent for this particular task.
Điều này sẽ thất bại vì chúng tôi đang sử dụng một tác nhân phụ cho nhiệm vụ cụ thể này.
The sub agent is actually getting the task right, but there's a communication breakdown between our sub agent and our orchestrator.
Tác nhân phụ thực sự đang thực hiện đúng nhiệm vụ, nhưng có sự cố giao tiếp giữa tác nhân phụ và bộ điều phối của chúng tôi.
This is a really common point of failure that we see when
Đây là một điểm thất bại rất phổ biến mà chúng tôi thấy khi
customers have really complicated systems with a lot of sub agents.
khách hàng có các hệ thống thực sự phức tạp với nhiều tác nhân phụ.
It's important to get the communication between your sub agents and your orchestrator just right.
Điều quan trọng là phải thiết lập giao tiếp giữa các tác nhân phụ và bộ điều phối của bạn một cách chính xác.
In the case of F2, like you see on the screen, this is an eval that's going to fail because we have a breakdown in that communication.
Trong trường hợp của F2, như bạn thấy trên màn hình, đây là một bài đánh giá sẽ thất bại vì chúng tôi có sự cố trong giao tiếp đó.
The last one that I'll highlight that you'll see fails is R8 on the screen.
Bài đánh giá cuối cùng mà tôi sẽ nhấn mạnh và bạn sẽ thấy thất bại là R8 trên màn hình.
R8 will essentially check the forecasting during a particular promotion month.
R8 về cơ bản sẽ kiểm tra dự báo trong một tháng khuyến mãi cụ thể.
This eval is also going to fail because we have two different policies that live in very different parts of our system prompt and actually end up contradicting each other.
Bài đánh giá này cũng sẽ thất bại vì chúng tôi có hai chính sách khác nhau nằm ở những phần rất khác nhau trong lời nhắc hệ thống của chúng tôi và cuối cùng mâu thuẫn với nhau.
So I mentioned over time our system prompt has grown.
Vì vậy, tôi đã đề cập rằng theo thời gian, lời nhắc hệ thống của chúng tôi đã phát triển.
We start to have some conflicts and the model gets confused leading towards a failure for this particular eval.
Chúng tôi bắt đầu có một số xung đột và mô hình bị nhầm lẫn dẫn đến thất bại cho bài đánh giá cụ thể này.
Now in the repo, you'll see it in the readme when we run
Bây giờ trong kho lưu trữ, bạn sẽ thấy nó trong tệp readme khi chúng tôi chạy
these evals, you'll see that they're going to pass up front at about 83% which is okay, but if you work in the world of manufacturing, that is not okay.
các bài đánh giá này, bạn sẽ thấy chúng sẽ vượt qua ban đầu ở mức khoảng 83%, điều này ổn, nhưng nếu bạn làm việc trong lĩnh vực sản xuất, điều đó không ổn.
17% failure is a really expensive failure percentage.
Tỷ lệ thất bại 17% là một tỷ lệ thất bại rất đắt đỏ.
Now let's double click on R8 again just so that we can understand a little bit about what's happening behind the scenes.
Bây giờ chúng ta hãy xem xét kỹ hơn R8 một lần nữa để hiểu một chút về những gì đang diễn ra đằng sau hậu trường.
Again, R8 is where we're essentially calculating the forecast during a particular month with a promotion.
Một lần nữa, R8 là nơi chúng tôi về cơ bản tính toán dự báo trong một tháng cụ thể với khuyến mãi.
And so in my on my screen here on the right side where you see kind of the simulated terminal window within the first block under the
Và vì vậy trên màn hình của tôi ở phía bên phải, nơi bạn thấy loại cửa sổ terminal mô phỏng trong khối đầu tiên dưới
commented text we can see that the agent pulled the right forecasting baseline and also pulled the right promotion multiplier.
văn bản đã được chú thích, chúng ta có thể thấy rằng tác nhân đã lấy đúng đường cơ sở dự báo và cũng lấy đúng hệ số nhân khuyến mãi.
So forecasting baseline 12 units a day promotion multiplier 3.1x this is all correct but in the calculation part below that we can see that there was actually some kind of hallucination that happened instead of using that 3.1x promo multiplier the
Vì vậy, đường cơ sở dự báo 12 đơn vị mỗi ngày, hệ số nhân khuyến mãi 3.1x, tất cả đều đúng, nhưng trong phần tính toán bên dưới, chúng ta có thể thấy rằng thực sự đã xảy ra một số loại ảo giác, thay vì sử dụng hệ số nhân khuyến mãi 3.1x đó,
agent actually ended up using 1.35.
tác nhân thực sự đã sử dụng 1.35.
So something happened along the way.
Vì vậy, điều gì đó đã xảy ra trên đường đi.
A hint here is that the reason for this is that we have context problems.
Một gợi ý ở đây là lý do cho điều này là chúng tôi có vấn đề về ngữ cảnh.
So this isn't a model problem.
Vì vậy, đây không phải là vấn đề của mô hình.
It's an issue with our the information that we're surrounding the model with.
Đó là vấn đề với thông tin mà chúng tôi đang bao quanh mô hình.
Our system prompt has grown to be really long and is very confusing for the model and has some conflicts in it which lead to the issue
Lời nhắc hệ thống của chúng tôi đã trở nên rất dài và rất khó hiểu đối với mô hình và có một số xung đột trong đó dẫn đến vấn đề
that shows up within this eval.
xuất hiện trong bài đánh giá này.
So folks, our objective in this workshop will first be to run our suite of eval.
Vì vậy, các bạn, mục tiêu của chúng tôi trong hội thảo này trước tiên sẽ là chạy bộ bài đánh giá của chúng tôi.
We're going to triage the issues and we're going to update the design of our agent accordingly.
Chúng tôi sẽ phân loại các vấn đề và cập nhật thiết kế của tác nhân cho phù hợp.
And then we're going to do something that we call internally hill climbing towards eval improvement.
Và sau đó chúng tôi sẽ làm điều mà chúng tôi gọi nội bộ là leo đồi hướng tới cải thiện bài đánh giá.
Right?
Đúng không?
So we run our evals, we get a baseline.
Vì vậy, chúng tôi chạy các bài đánh giá, chúng tôi có một đường cơ sở.
It's going to be about 83%.
Nó sẽ vào khoảng 83%.
We're then going to optimize the architecture of our agent and we're going to continue then running our eval so that we climb on them hopefully seeing the success percentage improve over time.
Sau đó, chúng tôi sẽ tối ưu hóa kiến trúc của tác nhân và tiếp tục chạy bài đánh giá để chúng tôi leo lên chúng, hy vọng thấy tỷ lệ thành công cải thiện theo thời gian.
In this lab, we're also going to start with an agent that is self-created on our messages API.
Trong phòng thí nghiệm này, chúng tôi cũng sẽ bắt đầu với một tác nhân được tự tạo trên API messages của chúng tôi.
Again, if you have the repo and you click on the before folder, I'll show
Một lần nữa, nếu bạn có kho lưu trữ và nhấp vào thư mục before, tôi sẽ chỉ
you this in just a bit.
cho bạn điều này trong một chút.
This is an agent that is built from scratch on our messages API.
Đây là một tác nhân được xây dựng từ đầu trên API messages của chúng tôi.
We're going to actually migrate that agent to Claude managed agents.
Chúng tôi sẽ thực sự di chuyển tác nhân đó sang các tác nhân được quản lý bởi Claude.
Claude managed agents essentially allows us to offload the messiness that comes with maintaining an agentic harness in scaling agents safely and securely to thousands and tens of thousands of users, right?
Các tác nhân được quản lý bởi Claude về cơ bản cho phép chúng tôi giảm tải sự lộn xộn đi kèm với việc duy trì một bộ khung tác nhân và mở rộng quy mô tác nhân một cách an toàn và bảo mật cho hàng nghìn và hàng chục nghìn người dùng, đúng không?
Like if I want to build my agent locally and run it locally, I can do that pretty quickly and pretty easily.
Giống như nếu tôi muốn xây dựng tác nhân của mình cục bộ và chạy nó cục bộ, tôi có thể làm điều đó khá nhanh và khá dễ dàng.
But the moment that I need to take that agent, I need to host it
Nhưng ngay khi tôi cần đưa tác nhân đó, tôi cần lưu trữ nó
remotely and I need to allow hundreds and thousands of users to at the same time engage with that agent.
từ xa và tôi cần cho phép hàng trăm và hàng nghìn người dùng cùng lúc tương tác với tác nhân đó.
There's an infrastructure problem, there's a scaling problem, there's memory, there's security, there's so much that I have to account for.
Có vấn đề về cơ sở hạ tầng, vấn đề mở rộng quy mô, bộ nhớ, bảo mật, rất nhiều thứ tôi phải tính đến.
So in order to offload that so I can just worry about the architecture of my agent itself and make decisions around tools, skills and sub agents.
Vì vậy, để giảm tải điều đó để tôi chỉ có thể lo lắng về kiến trúc của chính tác nhân và đưa ra quyết định về các công cụ, kỹ năng và tác nhân phụ.
I'm going to offload everything else to Claude managed agents.
Tôi sẽ giảm tải mọi thứ khác cho các tác nhân được quản lý bởi Claude.
So again, to break that
Vì vậy, một lần nữa, để phân tích
down just a bit, there's been a few talks on CMA so far today, but this is really where we're able to separate the agent from the session details from the sandboxed environment where tool calls are actually happening.
một chút, đã có một vài bài nói về CMA hôm nay, nhưng đây thực sự là nơi chúng tôi có thể tách tác nhân khỏi chi tiết phiên khỏi môi trường sandbox nơi các cuộc gọi công cụ thực sự diễn ra.
Again, this allows us to offload particular parts of the stack to then worry about the to then only worry about the design of our agent itself.
Một lần nữa, điều này cho phép chúng ta tách riêng các phần cụ thể của stack để chỉ tập trung vào thiết kế của agent.
All right.
Được rồi.
I mentioned that we're going to get hands-on in this workshop.
Tôi đã đề cập rằng chúng ta sẽ thực hành trong workshop này.
We are going to go ahead and do that right now.
Chúng ta sẽ bắt tay vào làm ngay bây giờ.
Now, what you see on the screen here is the workshop URL as well.
Trên màn hình bây giờ là URL của workshop.
If you haven't had a chance to grab it, feel free to go ahead and do so.
Nếu bạn chưa kịp lấy, hãy thoải mái làm điều đó.
This is where we're keeping all of the different workshops throughout Code with Claude within London, so you can go back and revisit them if helpful.
Đây là nơi chúng tôi lưu trữ tất cả các workshop khác nhau trong Code with Claude tại London, để bạn có thể quay lại xem lại nếu cần.
Within this workshop, we're going to be working on agent decomposition.
Trong workshop này, chúng ta sẽ làm việc về phân rã agent.
So that's going to
Đó sẽ là
be the name of the folder that we're actually going to be working within.
tên của thư mục mà chúng ta sẽ làm việc.
Great.
Tuyệt vời.
Let me jump forward here.
Để tôi chuyển tiếp ở đây.
Perfect.
Hoàn hảo.
So the first thing that we're going to do as a part of this workshop is we're first going to get a baseline.
Vậy điều đầu tiên chúng ta làm trong workshop này là thiết lập một baseline.
So when you open up that link, you'll first clone the repo.
Khi bạn mở link đó, trước tiên bạn sẽ clone repo.
So, we're going to clone the repo locally.
Chúng ta sẽ clone repo về máy local.
We have a UV project that's set up.
Chúng tôi đã thiết lập một dự án UV.
So, we're going to run UV sync in order to make sure that we have all of our packages and our dependencies
Chúng ta sẽ chạy UV sync để đảm bảo có tất cả các gói và phụ thuộc
to be able to invoke the Anthropic SDK and then eventually deploy our agent to Claude managed agents.
để có thể gọi Anthropic SDK và cuối cùng triển khai agent lên Claude managed agents.
So, we can run UV sync to do that.
Vậy chúng ta có thể chạy UV sync để làm điều đó.
I mentioned previously that we're going to need an API key for this workshop as well.
Tôi đã đề cập trước đó rằng chúng ta cũng cần một API key cho workshop này.
So using those credits that you got at the start of this session, you can go to your Claude console account and create an API key.
Sử dụng credits bạn nhận được đầu buổi, bạn có thể vào tài khoản Claude console và tạo API key.
If you copy the ENV example, you'll
Nếu bạn copy file ENV example, bạn sẽ
just have to manually copy your API key into the ENV file that's created for you.
chỉ cần thủ công copy API key vào file ENV được tạo cho bạn.
Now, all the 12 evals that I previously walked you through, we have all of those set up already.
Tất cả 12 bài eval mà tôi đã trình bày trước đó đều đã được thiết lập sẵn.
So in order to get a baseline and run those evals, you have to run uv run eval-agent before.
Để có baseline và chạy các eval đó, bạn cần chạy uv run eval-agent trước.
This is all in the readme, but if you just run that command, you will be able to actually go about running your evals.
Tất cả đều có trong readme, nhưng nếu bạn chạy lệnh đó, bạn sẽ có thể chạy các eval của mình.
Now, in terms of our building here, we're going to take a number of steps to actually go about running our evals, using Claude Code to triage the results of them, and then climbing accordingly on our agent.
Về phần xây dựng, chúng ta sẽ thực hiện nhiều bước để chạy eval, dùng Claude Code để phân loại kết quả và cải thiện agent tương ứng.
So, we're first going to take a look at our the system prompt that we have for our agent itself.
Đầu tiên, chúng ta sẽ xem xét system prompt của agent.
So, I mentioned earlier that our system prompt is
Tôi đã đề cập trước đó rằng system prompt của chúng ta
currently sitting at about 400 lines long.
hiện tại dài khoảng 400 dòng.
We've been stacking information on our system prompt over and over again as we've continued to get more business requirements.
Chúng ta đã liên tục chồng chất thông tin lên system prompt khi có thêm yêu cầu kinh doanh.
So our system prompt is very long.
Vì vậy system prompt rất dài.
We'll take a look at that.
Chúng ta sẽ xem xét nó.
We are then going to take some time to evaluate the tools that we're using.
Sau đó, chúng ta sẽ dành thời gian đánh giá các công cụ đang sử dụng.
Right now, as I mentioned, we have 12 different tools.
Hiện tại, như tôi đã đề cập, chúng ta có 12 công cụ khác nhau.
Three of them are actually kind of wrapped sub agents.
Ba trong số đó thực chất là các sub agent được bọc lại.
So, we'll take a look to see what we can do to make that more efficient.
Chúng ta sẽ xem xét cách làm cho hiệu quả hơn.
And then lastly, if there are any sub agents that we really need to make our agent effective, we're going to take a look at the best way to actually construct sub agents with Claude managed agents.
Và cuối cùng, nếu có sub agent nào thực sự cần thiết để agent hoạt động hiệu quả, chúng ta sẽ xem xét cách tốt nhất để xây dựng sub agent với Claude managed agents.
I'm going to jump back just for a moment.
Tôi sẽ quay lại một chút.
There's one thing that I forgot to mention for you as you get started.
Có một điều tôi quên đề cập khi bạn bắt đầu.
Within the repo folder, there's two different folders that you'll see.
Trong thư mục repo, bạn sẽ thấy hai thư mục khác nhau.
There's a before folder and then there's a starter folder.
Có thư mục before và thư mục starter.
Those contain two separate agents.
Chúng chứa hai agent riêng biệt.
So, if you want to view the messages API version of the agent, again, this is just me building my own agent loop and my own agent harness around the Anthropic messages API to invoke Claude.
Nếu bạn muốn xem phiên bản messages API của agent, đây là tôi tự xây dựng vòng lặp agent và harness quanh Anthropic messages API để gọi Claude.
You'll see that within the before folder.
Bạn sẽ thấy nó trong thư mục before.
If you want to view what that agent looks like when deployed on Claude managed agents you can look in the starter folder which exists right below that.
Nếu bạn muốn xem agent đó khi triển khai trên Claude managed agents, bạn có thể xem trong thư mục starter nằm ngay bên dưới.
If you want to deploy your agent on Claude managed agents you can run uv run deploy starter.
Nếu bạn muốn triển khai agent lên Claude managed agents, bạn có thể chạy uv run deploy starter.
So again run your evals using the messages API version - agent before you can then deploy your agent on Claude managed agents.
Vậy hãy chạy eval của bạn bằng phiên bản messages API - agent before, sau đó bạn có thể triển khai agent lên Claude managed agents.
We
Chúng tôi