{"id":304,"date":"2026-09-07T14:28:22","date_gmt":"2026-09-07T08:58:22","guid":{"rendered":"https:\/\/tierones.io\/blog\/?p=304"},"modified":"2026-09-07T14:28:24","modified_gmt":"2026-09-07T08:58:24","slug":"intern-evaluation-decision-rubric","status":"publish","type":"post","link":"https:\/\/tierones.io\/blog\/intern-evaluation-decision-rubric\/","title":{"rendered":"62% of a Performance Rating Is the Rater, Not the Person. Write the Intern&#8217;s Rubric Before Week One."},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">When Scullen, Mount and Goff decomposed supervisory performance ratings across two samples of managers \u2014 2,350 and 2,142 people, each rated by seven raters \u2014 <a href=\"https:\/\/www.semanticscholar.org\/paper\/Understanding-the-latent-structure-of-job-ratings.-Scullen-Mount\/0a73fb7d291a407656d4ee4a9b1eb19514abe157\" rel=\"nofollow noopener\" target=\"_blank\">idiosyncratic rater effects accounted for 62% and 53% of the variance in the ratings, while actual ratee performance accounted for 21% and 25%<\/a>. Their paper ran in the <em>Journal of Applied Psychology<\/em> in 2000 and the finding has aged well. Read it plainly: most of what a rating measures is the person holding the pen. That is the instrument most startups use to decide whether an intern gets an offer.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Key takeaways<\/h2>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Unstructured end-of-internship impressions mostly measure the evaluator, not the intern.<\/li>\n\n\n\n<li>Criteria invented after you have formed a view get bent to fit that view \u2014 the effect disappears when you commit to them first.<\/li>\n\n\n\n<li>Entry-level hiring at early-stage startups is down about 76% against 2019, which makes each intern decision a larger share of your junior pipeline than it used to be.<\/li>\n\n\n\n<li>Decide during the final week, not after it. Conversion is at a five-year high and 88.3% of offers are being accepted.<\/li>\n<\/ul>\n\n\n\n<h2 class=\"wp-block-heading\">The rating is mostly about you<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The Scullen result is not a claim that managers are careless. It is a statistical decomposition: when the same person is rated by several people on several dimensions, the largest consistent source of variance is which rater is doing the rating. One founder runs generous; another anchors on a single bad week. Both feel equally certain.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For a twelve-person startup this is worse, not better, than it is at a large company, because there is usually exactly one rater. There is no calibration meeting, no second reviewer, no distribution to compare against. Whatever correction exists has to be built into the process deliberately, because the org chart will not supply it.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Criteria written afterwards are not criteria<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The second failure is subtler and more expensive. In <a href=\"https:\/\/pubmed.ncbi.nlm.nih.gov\/15943674\/\" rel=\"nofollow noopener\" target=\"_blank\">Uhlmann and Cohen&#8217;s 2005 studies in <em>Psychological Science<\/em><\/a>, participants choosing between job candidates did not report that the candidates had different strengths \u2014 they redefined what the job required so that the requirement matched whichever candidate they already preferred. Commitment to the hiring criteria before the applicant&#8217;s details were disclosed eliminated the effect.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">An internship is a long, information-rich version of exactly that setup. You form an impression in week two. By week six the standard has quietly rearranged itself: if your favourite ships fast and communicates poorly, the job becomes about shipping fast; if they write beautifully and ship slowly, the job becomes about judgment. Both stories are available, which is the problem. The only defence is that the criteria were fixed while you still did not know who would satisfy them.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"576\" src=\"https:\/\/tierones.io\/blog\/wp-content\/uploads\/2026\/09\/rubric-before-evidence-inline-opt-1024x576.jpg\" alt=\"rubric-before-evidence-inline\" class=\"wp-image-310\" srcset=\"https:\/\/tierones.io\/blog\/wp-content\/uploads\/2026\/09\/rubric-before-evidence-inline-opt-1024x576.jpg 1024w, https:\/\/tierones.io\/blog\/wp-content\/uploads\/2026\/09\/rubric-before-evidence-inline-opt-300x169.jpg 300w, https:\/\/tierones.io\/blog\/wp-content\/uploads\/2026\/09\/rubric-before-evidence-inline-opt-768x432.jpg 768w, https:\/\/tierones.io\/blog\/wp-content\/uploads\/2026\/09\/rubric-before-evidence-inline-opt-1536x864.jpg 1536w, https:\/\/tierones.io\/blog\/wp-content\/uploads\/2026\/09\/rubric-before-evidence-inline-opt.jpg 1600w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<h2 class=\"wp-block-heading\">The decision carries more weight than it used to<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">SignalFire&#8217;s State of Talent report, published in June 2026 from its Beacon dataset covering more than 650 million individuals, found new-graduate hiring <a href=\"https:\/\/www.signalfire.com\/blog\/signalfire-state-of-talent-report-2026\" rel=\"nofollow noopener\" target=\"_blank\">down roughly 76% at early-stage startups and around 65% at the tech majors against 2019<\/a>. Whatever you believe about why, the arithmetic is unambiguous: fewer junior hires are being made, so the ones you do make are a much larger fraction of your future senior bench.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Meanwhile the market for the good ones has tightened. In NACE&#8217;s 2026 Internship &amp; Co-op Survey of 284 organisations, run between October 2025 and January 2026, <a href=\"https:\/\/career.lafollette.wisc.edu\/blog\/2026\/04\/08\/intern-conversion-rate-hits-highest-mark-in-five-years\/\" rel=\"nofollow noopener\" target=\"_blank\">the conversion rate hit 63.1% \u2014 a five-year high \u2014 while acceptances rose to 88.3% from 82.8%<\/a>. Employers are converting more interns and interns are saying yes more often. A decision you defer for three weeks after the last day is a decision someone else has already made for you.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">The protocol<\/h2>\n\n\n\n<ol class=\"wp-block-list\">\n<li><strong>Week 0 \u2014 write four to six criteria and freeze them.<\/strong> Concrete and observable: ships a scoped change without a second reminder; writes a pull request another engineer can review cold; asks before burning a day on an ambiguous ticket; recovers from a broken build without escalation.<\/li>\n\n\n\n<li><strong>Say them out loud on day one.<\/strong> An intern who knows the standard is being measured against it. One who does not is being surprised by it.<\/li>\n\n\n\n<li><strong>Log evidence weekly, not opinions.<\/strong> Two lines a week with a link \u2014 a merged PR, a thread, an incident. Ten minutes, and it is the only thing standing between you and a memory reconstructed from the last fortnight.<\/li>\n\n\n\n<li><strong>Run a mid-point review against the frozen list.<\/strong> Halfway through, score it and say what would have to change. A criterion nobody can find evidence for was probably never the job.<\/li>\n\n\n\n<li><strong>Get a second rater on the artefacts.<\/strong> Have another engineer read two pull requests and one written summary without your commentary attached. It is the closest thing a small team has to calibration.<\/li>\n\n\n\n<li><strong>Score before you discuss.<\/strong> Everyone writes their scores down independently, then you compare. Discussion first converges the room on whoever speaks first.<\/li>\n\n\n\n<li><strong>Decide in the final week and say it plainly.<\/strong> Yes with a start date and terms; or no, with the two specific things that were missing. A vague no costs you the referral and the second look a year from now.<\/li>\n<\/ol>\n\n\n\n<h2 class=\"wp-block-heading\">This is the same discipline that made the trial worth running<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">A scoped work trial only produces signal if the reading of it is as structured as the setting of it \u2014 otherwise you have spent six weeks generating evidence and then ignored it. We have written before on <a href=\"https:\/\/tierones.io\/blog\/six-week-internship-work-trial\/\">running a six-week internship as a work trial<\/a> and on <a href=\"https:\/\/tierones.io\/blog\/convert-remote-interns-full-time\/\">converting remote interns into full-time hires<\/a>; this post is the missing middle between them, the part where the evidence becomes a decision.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Tierones works the same way at the front of the funnel. Students from IITs, NITs and IIITs come to you with proof of work \u2014 commits, shipped projects, a recorded walkthrough of code they wrote \u2014 so the first judgement you make is about evidence rather than a r\u00e9sum\u00e9. Stipends are set by you, not by us. What we can do is make sure the thing you are evaluating is real. <a href=\"https:\/\/tierones.io\/employers\">Tell us what you need<\/a> and we&#8217;ll show you proof, not CVs.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">FAQ<\/h2>\n\n\n\n<h3 class=\"wp-block-heading\">How should a small startup evaluate an intern at the end of a placement?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Against criteria written before the placement began, using evidence logged while it ran. Scullen, Mount and Goff found idiosyncratic rater effects accounting for 62% and 53% of rating variance against 21% and 25% for actual performance. A frozen rubric and dated links are the cheapest correction available to a team of twelve.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Why does it matter whether the criteria are written down in advance?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Because criteria written afterwards get shaped by the conclusion. Uhlmann and Cohen found participants redefining what a job required to match the candidate they already preferred, and the effect disappeared when they committed to the criteria first. Over six weeks with one intern, the same drift is almost invisible from the inside.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">How quickly should we make the decision after the internship ends?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Before it ends. NACE&#8217;s 2026 survey put intern conversion at 63.1%, a five-year high, with 88.3% of offers accepted. The strongest candidates are being converted quickly, so a decision that drifts into the weeks after the last day is usually a decision made by a competitor.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><em>Hiring remote engineering help from Indian Tier-1 campuses and want the evaluation to start from evidence? <a href=\"https:\/\/tierones.io\/employers\">Tell us what you need<\/a> \u2014 we&#8217;ll show you proof, not CVs. Questions go to <a href=\"mailto:hire@tierones.io\">hire@tierones.io<\/a>.<\/em><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Research on 4,492 managers found 62% of rating variance came from the rater and 21% from actual performance. Here is how to run the end-of-internship call.<\/p>\n","protected":false},"author":1,"featured_media":309,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[3],"tags":[66,198,197,57,199,137,200,43],"class_list":["post-304","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-for-employers","tag-engineering-management","tag-hiring-rubric","tag-intern-performance-evaluation","tag-remote-interns","tag-return-offer","tag-startup-hiring","tag-structured-evaluation","tag-us-startups"],"_links":{"self":[{"href":"https:\/\/tierones.io\/blog\/wp-json\/wp\/v2\/posts\/304","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/tierones.io\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/tierones.io\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/tierones.io\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/tierones.io\/blog\/wp-json\/wp\/v2\/comments?post=304"}],"version-history":[{"count":1,"href":"https:\/\/tierones.io\/blog\/wp-json\/wp\/v2\/posts\/304\/revisions"}],"predecessor-version":[{"id":311,"href":"https:\/\/tierones.io\/blog\/wp-json\/wp\/v2\/posts\/304\/revisions\/311"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/tierones.io\/blog\/wp-json\/wp\/v2\/media\/309"}],"wp:attachment":[{"href":"https:\/\/tierones.io\/blog\/wp-json\/wp\/v2\/media?parent=304"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/tierones.io\/blog\/wp-json\/wp\/v2\/categories?post=304"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/tierones.io\/blog\/wp-json\/wp\/v2\/tags?post=304"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}