You've got twelve resumes open, three “performance marketing ninjas” waiting for interviews, and one very real problem: none of them has shown you whether they can manage a paid media account when the numbers turn ugly.
The interviews go well. Everyone speaks confidently about ROAS, funnels, attribution, and scaling. Then the new hire touches the account, costs rise, creative testing becomes a guessing game, and the dashboard starts looking like a crime scene.
That's what happens when candidate evaluation criteria describe a person instead of measuring performance. “Strategic.” “Data-driven.” “Great communicator.” Lovely words. Completely useless unless you can observe what those traits produce inside an actual advertising workflow.
A serious evaluation system should reduce noise, increase signal, and let two interviewers reach comparable conclusions about the same candidate. It should also expose résumé theater, polished AI-generated answers, and impressive claims that collapse under one follow-up question.
The evidence supports structure. Research has consistently found that structured interviews predict job performance better than unstructured conversations. The 1998 Schmidt and Hunter meta-analysis reported corrected validity of 0.51 for structured interviews versus 0.38 for unstructured interviews, making structured interviews about 34% more predictive in that analysis (historical comparison of structured and unstructured interviews).
For media buyers, though, interviews shouldn't carry the whole burden. The role produces observable work. Your evaluation should inspect that work directly.
Hiring a media buyer without defined criteria usually starts innocently. You skim a résumé, notice familiar platform names, and feel relieved when the candidate says they've “scaled multiple accounts.” In the interview, they're quick, articulate, and confident. You like them.
Then you ask what they changed.
The answer gets foggy. They mention a strong client result but can't explain the baseline, attribution model, offer, budget conditions, creative volume, or what they personally owned. They describe “optimizing campaigns” without naming the decision rule. They say they improved ROAS, but not whether revenue quality changed or whether the account benefited from a seasonal spike.
That isn't evidence. It's advertising folklore.
Practical rule: If a candidate can't explain the decision behind a result, don't score the result as proof of skill.
Candidate evaluation criteria should function like a performance prediction system. They should tell you what matters, how you'll observe it, and what separates acceptable execution from excellent execution. A résumé can suggest exposure. A structured interview can reveal reasoning. A work sample can show whether the candidate can do the job without a rescue helicopter.
The strongest system is deliberately narrow. It doesn't reward candidates for collecting every platform badge or reciting every marketing acronym. It measures the few capabilities that determine whether paid media work survives contact with a real budget.
That means assessing how someone diagnoses an account, chooses a testing priority, allocates spend, interprets incomplete data, communicates tradeoffs, and learns from failure. It also means making every candidate face comparable prompts and scoring their answers against anchored standards.
A founder doesn't need another hiring ritual. You need a repeatable way to answer one question: Will this person make better paid media decisions than the person we have now?
The rest is résumé confetti.
Credentials tell you what a candidate has touched. Traits tell you how they describe themselves. Job-linked signals show what they can produce. Those categories overlap, but they aren't interchangeable.
Hiring a pilot because they wrote a thoughtful essay about flying would be absurd. You'd want to see how they handle the aircraft. Media buyer hiring deserves the same level of common sense. Platform experience matters, but the useful question is whether the candidate can turn platform access into sound decisions under commercial constraints.
Three types of validity make that distinction practical.
If the role involves building campaigns, analyzing performance, and deciding where the next dollar goes, your assessment should contain those activities. Asking someone to explain a media-buying philosophy has weaker content validity than giving them a messy account snapshot and asking what they'd investigate first.
The closer the task mirrors the role, the less you rely on proxies such as degrees, employer logos, or polished language. The candidate has to demonstrate the work instead of performing competence theatrically.
A criterion is the outcome you care about, such as sound optimization decisions, reliable reporting, or effective campaign execution. Criterion-related validity asks whether an assessment score relates to later job performance.
The McDaniel et al. meta-analysis reviewed 106 structured-interview studies involving 12,847 candidates. It found mean operational validity of 0.44 for structured interviews versus 0.33 for unstructured interviews, while structured board interviews using consensus ratings reached as high as 0.64 in corrected analyses (McDaniel et al. criterion-related validity research).
The lesson isn't that a score predicts everything. It's that standardization, multiple raters, and consistent scoring improve the quality of the signal.
A résumé screen and an interview may tell you something. A work sample can add information they can't provide. That extra signal matters when candidates have similar backgrounds or when credentials are easy to inflate.
The U.S. Office of Personnel Management describes structured interviews as having high content validity, high criterion-related validity, and incremental validity, because they can add useful signal beyond other assessment methods (OPM guidance on structured interviews).
For paid media, the practical formula is simple: use the résumé to establish context, the interview to inspect reasoning, and the work sample to inspect execution.

A scorecard should contain five sharp criteria, not a museum of every desirable human quality. Start with the work your business needs, then define what acceptable and exceptional performance look like.
This covers platform mechanics across the channels the role will manage, such as Meta Ads, Google Ads, TikTok Ads, LinkedIn, or Microsoft Ads. Look for campaign architecture, tracking awareness, audience setup, bidding logic, creative formats, and troubleshooting ability.
A strong candidate doesn't merely list platforms. They can explain why a campaign was structured a certain way and what they'd change when delivery, tracking, or audience quality breaks.
Strategy appears in tradeoffs. Can the candidate connect offer, funnel stage, audience, creative angle, budget, and conversion event? Can they decide whether to scale, consolidate, refresh creative, repair the landing page, or stop spending?
A media buyer who treats every problem as a bid adjustment is not strategic. They're operating a very expensive keyboard.
Look for the ability to separate symptoms from causes. The candidate should interpret metrics in context, question attribution, identify data limitations, and explain what evidence would change their decision.
Strong analysts don't worship dashboards. They use them to form hypotheses, test them, and communicate uncertainty without hiding behind it.
Paid media performance depends on more than audience settings. Assess how the candidate develops angles, briefs creative, evaluates fatigue, compares tests, and decides when a concept has earned more spend.
Ask for a testing roadmap rather than a list of ad formats. The roadmap reveals whether they understand learning, prioritization, and the relationship between creative and funnel behavior.
An excellent buyer can build clean account structures, maintain a reliable optimization cadence, document decisions, and report clearly. They can also explain bad news to a client or internal team without turning the meeting into interpretive dance.
Weight these criteria according to the business model. A DTC brand may place more emphasis on creative iteration and unit economics. B2B SaaS may require stronger lead-quality judgment, funnel analysis, and CRM feedback loops. An agency may need heavier weighting on communication, prioritization, and managing multiple accounts.
Write the scorecard before reviewing candidates. The media buyer job description guide can help translate responsibilities into observable evaluation areas, but don't copy a generic checklist and call it strategy.

A scorecard becomes useful only when the numbers mean something. A bare 1-to-5 scale creates the illusion of rigor while allowing every interviewer to grade from personal instinct.
Use anchored scoring. For each criterion, define what weak, competent, strong, and exceptional evidence looks like. A candidate who earns a high score in analytical rigor should demonstrate a repeatable way to investigate performance, not merely use the word “data” several times.
Here's a sample scorecard for a hands-on media buyer. Adjust the weights to the role, but keep the criteria narrow enough that one person can remember what each score represents.
| Criterion | Weight | What 5 Looks Like | Scoring Prompt |
|---|---|---|---|
| Technical mastery | 25% | Builds and troubleshoots channel setups with clear platform logic | Can the candidate explain the setup, diagnose delivery issues, and defend implementation choices? |
| Strategic thinking | 25% | Connects business goals to funnel, audience, offer, budget, and testing priorities | Does the candidate choose actions based on commercial impact rather than habit? |
| Analytical rigor | 20% | Forms clear hypotheses, identifies data limits, and makes evidence-based decisions | Can the candidate distinguish correlation, attribution noise, and actionable signal? |
| Creative testing judgment | 15% | Designs useful tests and knows when to iterate, scale, or stop | Does the candidate turn performance data into better creative decisions? |
| Operational discipline and partnership | 15% | Documents work, communicates tradeoffs, and maintains dependable execution | Would clients and teammates know what is happening and why? |
Calculate each criterion's score, multiply it by its weight, and add the weighted results. A candidate scoring 4 in technical mastery contributes 1.00 to the composite score under this model. A 3 in strategic thinking contributes 0.75. The arithmetic isn't the magic. The discipline is.
Structured interviews are useful because they standardize questions, ordering, and scoring. Research summaries place structured interview validity around 0.42 to 0.51, compared with roughly 0.19 to 0.38 for unstructured interviews (summary of structured interview research).
Work samples test the work itself. Meta-analytic evidence places their corrected validity around 0.33 for overall job performance, and as high as 0.54 for current-job performance in some analyses (Roth et al. work-sample research).
Have interviewers score independently before the debrief. Otherwise, the most senior voice sets the mood, and everyone else starts grading the candidate's charisma instead of the evidence. A skills-based hiring framework gives teams a useful foundation, but the final rubric still needs to reflect your account, customers, and constraints.
The best assessment workflow combines conversation with execution. Start with a structured interview, then give the candidate a compact task that resembles the job. Don't assign a week-long unpaid campaign rebuild. You're hiring a buyer, not commissioning a free consulting project.
Use the same questions for every candidate and map each question to one criterion.
Follow-up questions should test ownership. Ask what the candidate personally changed, what failed, what they'd do differently, and which data they didn't have. Vague answers become obvious when the interviewer keeps returning to decisions.
Account diagnosis: Provide a sanitized account snapshot with campaign structure, spend context, creative examples, and performance trends. Ask for a concise written diagnosis or a 48-hour Loom walkthrough covering the first issues they'd investigate, immediate risks, recommended actions, and missing information.
Testing roadmap: Provide raw CSV data and a short business brief. Ask the candidate to propose a testing roadmap, explain prioritization, define success conditions, and identify where the data might mislead them.
Score both samples with the same rubric used in the interview. That prevents the work sample from becoming a separate popularity contest.
A good test is hard to fake because it requires connected decisions, not trivia.
Give candidates enough context to work fairly, prohibit access to confidential data, and ask them to explain their reasoning live afterward. For remote candidates, request a screen-recorded walkthrough and include a short follow-up where they defend one recommendation. AI can polish language. It has a harder time defending a coherent decision tree under pressure.
Client references can validate ownership and communication, but they shouldn't replace direct testing. Use client reference checks for media buyers as one evidence source among several.

Paid media hiring has a particular weakness: candidates can borrow the language of performance without owning the decisions behind it. A résumé can be optimized. A case study can be polished. A dashboard screenshot can be stripped of every inconvenient detail.
Your job is to find the missing context.
AI-generated answers aren't automatically proof of dishonesty. They are a reason to test authorship and understanding. Ask the candidate to critique their own case study, recreate a recommendation from unfamiliar data, or explain why an alternative decision would have failed.
A strong buyer can say, “I killed this creative,” and explain why. They keep iteration logs, distinguish platform reporting from business outcomes, ask clarifying questions before prescribing tactics, and state what they'd need to know before making a confident recommendation.
They also understand that remote reliability isn't captured by competency scores alone. Check whether the candidate can work across time zones, document decisions, communicate blockers, and maintain a dependable operating rhythm. Long-term fit should be evaluated through observable working behavior, not a vague “culture fit” impression.
Years of experience don't guarantee judgment. Charisma doesn't guarantee client partnership. A famous brand on a résumé doesn't prove the candidate personally managed the account. More criteria don't create more accuracy, either. A narrower rubric tied to the job is usually more useful than a sprawling checklist designed to make everyone feel thorough.

A useful system fits into the hiring process instead of becoming another project nobody maintains. Use four stages.
After the hire, keep the scorecard alive. Review whether the criteria predicted early execution, communication quality, retention, and performance drift. If analytical rigor never separates strong hires from weak ones, either the anchor is poor or the criterion isn't doing the work you thought it was.
For teams that don't want to build sourcing and vetting from scratch, HireMediaBuyers.com offers a marketplace of pre-vetted media buyers and paid ads specialists, with technical assessments, interviews, skills testing, reference checks, and track-record verification described as part of its process. You still need your own role-specific scorecard, because no marketplace can know which tradeoffs matter most inside your business.
The checklist is straightforward: define five criteria, write behavioral anchors, weight business impact, standardize questions, test real work, score independently, and review outcomes after hiring. That's how candidate evaluation criteria become a management tool instead of an HR ornament.
If you need vetted media buyers or paid ads specialists to evaluate against a sharper, work-sample-driven scorecard, visit HireMediaBuyers.com. Browse talent or request a tailored shortlist, then judge candidates on the decisions they can make, not the marketing poetry on their résumés.