Judgment
Study judgment, not features: my CCAO-F plan as a Tech Lead
A four-week CCAO-F plan that practices judgment on synthetic client data before I sit the exam.
- Planted
- Last tended

I am a tech lead at an outsourcing company, so my domain changes with the project and the client. My reflex in every one of them is to build something: a connector, a script, a clever fix. The exam seems to reward judgment instead. I already open Claude for drafts, summaries, and the odd design note. That habit will not pass the Claude Certified Associate – Foundations exam. The exam guide spends more than half the score on checking output, fitting Claude into a workflow, and deciding what is safe to upload. I have not booked a seat, and I have not sat the exam. This is the four-week plan I will follow, dated as of 3 October 2026.
The score is a judgment score
The official name is Claude Certified Associate – Foundations (CCAO-F). The table is from the exam guide as of 3 October 2026. A multi-response item states how many options to select. Delivery is Pearson VUE, at a test center or OnVUE. List price is 99 USD. A registration stays valid for five years, and the credential lasts 12 months from the award.
| Fact | What I am planning around |
|---|---|
| Length | 60 items, 120 minutes |
| Pass | Scaled 720 out of 100–1000, plus domain percents |
| Heavy domains | Output evaluation 21%, workflow 16%, governance 15% (52% together) |
| Study spine | The free Associate prep course on Skilljar |
| Practice data | Synthetic or de-identified rows only. No live client extract |
| Result | None yet. A later post will carry the score if I sit |
The seven domains, and where the hours go
These weights are from the exam guide, as of 3 October 2026.
| Domain | Weight | What the item is really asking |
|---|---|---|
| Prompting and task execution | 14% | Can you structure the request and split a large task? |
| Output evaluation and validation | 21% | Do you check the reply before you put your name on it? |
| Product and model selection | 12% | Did you pick the surface and the model that fit this task? |
| Workflow integration | 16% | Which steps stay with a person, and which move to Claude? |
| Configuration and knowledge | 12% | Can a Project keep the right instructions and sources over time? |
| Governance, risk, and responsible use | 15% | Should this data, this claim, or this audience be in the chat at all? |
| Troubleshooting and optimization | 10% | Can you name why a reply failed and make the fix stick? |
Output evaluation, workflow integration, and governance are 52% of the exam. I am treating those three as the course. The other four are the vocabulary the judgment items assume.
The prep course is free to register. Skilljar does not require an Anthropic account to view the lessons. Using Claude is a separate account. I will take the eight lessons in order, and I will not treat the runtime as the whole plan:
- Claude Platform and Model Foundations, 59 minutes.
- Prompting and Task Execution, 53 minutes.
- Evaluating and Validating Claude's Output, 74 minutes.
- Workflow Integration and Solution Design, 63 minutes.
- Configuration and Knowledge Management, 47 minutes.
- Governance, Risk and Responsible Use, 55 minutes.
- Troubleshooting and Optimization, 30 minutes.
- Course Summary and Next Steps, 8 minutes.
That is about six and a half hours. Evaluation is the longest lesson and the heaviest domain. The page recommends Claude 101, AI Fluency: Framework and Foundations, and AI Capabilities and Limitations first, and it aims the course at knowledge workers in operations, marketing, project management, education, and communications. I will open those earlier courses only when a lesson uses a term I cannot define. My miss is the other direction: I reach for an integration when the item wants a check.
The trap I am drilling
The exam guide publishes a few sample items. I am paraphrasing the pattern, not copying the wording, and these are practice samples rather than live exam content.
One sample is a regulation summary that cites a subsection. The reply sounds finished. The move that holds up is to open the official text and read that subsection. Sending it because the prose is confident, or because Claude rated its own confidence, is trust. Rewriting it so it sounds more formal is polish. The citation is still unread.
Another sample is a pile of short drafts where speed and cost matter. The fitting move is a faster, cheaper model. The most capable model on every short reply is overkill: the impressive choice, on a task that did not ask for it.
A third sample is a spreadsheet of customer names and account numbers, under a policy that restricts regulated personal data. The move that holds up is to remove or anonymize the identifiers before upload. Uploading the file as it stands is trust. Skipping the analysis is giving up. Uploading it and adding "do not retain this" still puts the regulated file in the chat. That last one is the line I would have picked on a bad day, because it feels like a control.
The misses are four families: trust the model, polish the output, overkill, and give up. On a Tech Lead's reflex, overkill also includes a connector, a script, or a retention instruction when the Associate answer is to verify, anonymize, configure a Project, or escalate.
Press a family on the diagram. A wrong habit shakes. The source check pulses. The checklist on the right is the same rule in one line.
Interactive Demo
A reply cites a clause. Pick the habit you would act on before you forward it.
A purple box labeled Cited clause sits above four neutral replies and one teal reply labeled Open the source. Buttons name the four distractor families plus verify. Pressing a button draws a dashed line, moves a token, and shakes a wrong reply or pulses the source check.
- Trust the model sends the reply because it sounds sure.
- Polish rewrites the tone and still never opens the clause.
- Overkill spends a bigger model on a check a person can do.
- Give up drops the task instead of doing the one check.
- Verify opens the official text and reads the cited clause.
- Reset and try a family yourself.
Habit: not chosen
Code
Code
Read onlynotes/checklist.md
Read only
Before you forward a cited clause:
trust: a confident tone is not a check
polish: a formal rewrite is not a check
overkill: a larger model is not a check
giveup: skipping the task is not the task
verify: open the official text and read the clauseFour weeks on one synthetic set
Every drill uses rows I type myself, or a file already de-identified under the client's process and checked again by me. A live extract stays out unless I can name who approved it.
Week 1 — surface, model, and a Project. Lessons 1 and 2. I will write a Project whose instructions name the role, the output shape, and a hard line: do not invent a citation. The knowledge is a short style note and ten synthetic operational blurbs, with no names and no record numbers. This Project is the capstone. By Friday it has to exist and refuse a made-up clause number on a smoke test. The same evening I will run one short-draft task on two models and write which reply I would send, and why the other was the wrong spend.
Week 2 — the verification log. Lesson 3, the 74-minute one. Domain 2 is 21%, so this is the heavy week. For every factual claim I will log the claim, the source I opened, and pass or fail. One sitting uses a fake policy blurb that cites a subsection I know is wrong. I want to see whether I catch it before the log says pass. A confidence line from Claude is not a source. An empty source cell fails the item.
Week 3 — workflow map and de-identification. Lessons 4 and 5. I will draw one task I actually touch: an operational note into a status update for a lead. Each box is a person or Claude, and the handoff is a sentence. At least one step stays human because a wrong sentence would change a decision the client has to live with.
Then the data drill. I will type a fake row shaped like the current client: a made-up name, a made-up account or record number, a made-up date, and one field that the client treats as sensitive. The field changes with the project. I will delete identifiers until the remaining text could not pick out a person. Only that stripped text may enter the Project. The stripping checklist goes into Project knowledge so the next session starts from the rule.
Week 4 — break the Project, then the samples. Lessons 6 and 7, then the eight-minute summary. I will remove the instruction that forbids invented citations, watch the failure, put the instruction back, and write why the fix lives in the Project rather than in one chat. I will then redo the three sample patterns from memory and name each distractor's family before I look at the guide again.
What stays when the client changes
The spreadsheet sample is the same shape on every engagement. One client restricts account numbers. The next restricts a student ID, a ticket log, or a column I have not seen yet. I am not studying one domain's statute as if it were the job. The move is the job.
Identifiable client data stays off a third-party model. Drills use synthetic rows I type myself, or data already de-identified and checked again. If a real task needs a model, stripping comes first, and the vendor decision sits with the client's process. A prompt that says "do not retain this" is not that process.
Booking, once the drills exist
I will register with my Claude Partner Network organization email and pay the list price, 99 USD. The partner FAQ says a personal email will not carry the certification. The 50% rate and the 100% Global Premier window, which ended 31 August 2026, do not apply to this seat.
The Pearson name has to match my government photo ID exactly. I will enter Ngô Huy Hoàng. A passport is an accepted form. If that ID prints a Latin spelling without diacritics, I type that line instead: a mismatch means I cannot test, and the 99 USD is forfeited. A correction goes to [email protected] at least 24 hours ahead, with 24 to 48 business hours to process.
Pearson's Anthropic page, updated 2 October 2026, requires 48 hours to reschedule or cancel, or the fee is forfeited. The partner FAQ and the policies page still say 24 hours. I will live by 48 and re-read both pages the morning I book. Retakes wait 14 days, then 30, then 90, at most four attempts in a rolling 12 months, fee each time. An on-time renewal is a free non-proctored assessment on the Partner Academy. A lapse means the full exam again.
For OnVUE I will use a home network. The OnVUE page disallows VPN, corporate networks, and public networks, and it allows one screen. Check-in starts 30 minutes early. A digital whiteboard is allowed. The system test runs on the same machine and the same network.
What I will not fill in
I do not know whether this exam allows a break. Pearson says not every exam offers one, and the Anthropic allowance list I read does not name a break. I also do not know the late-arrival window, whether multi-response items award partial credit, or whether any items are unscored. Third-party pages disagree, so those four stay blank. A booking screen or a newer guide can fill them later. The plan does not depend on them.
After I sit
The follow-up, if I take the exam, will be the scaled score, the domain percents, and which drill predicted the weak domain. I will not write that post from a guess.
This note sits under AI Engineering, on the Claude & AI leaf, beside LLM Fundamentals and Agentic Workflows. Three notes should come out of the drills, and I have not written them yet: prompt anatomy for ordinary work tasks, a verification log for Claude's output, and stripping client identifiers before a row reaches a model.