{"id":177,"date":"2026-07-31T13:19:10","date_gmt":"2026-07-31T07:49:10","guid":{"rendered":"https:\/\/tierones.in\/blog\/?p=177"},"modified":"2026-07-31T13:19:11","modified_gmt":"2026-07-31T07:49:11","slug":"hiring-junior-developers-judgment-ai","status":"publish","type":"post","link":"https:\/\/tierones.io\/blog\/hiring-junior-developers-judgment-ai\/","title":{"rendered":"66% of Developers Say AI Output Is &#8220;Almost Right, But Not Quite.&#8221; That Gap Is the Job You&#8217;re Now Hiring For."},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">The single most common complaint from working developers in 2025 was not that AI is useless. It was that <a href=\"https:\/\/survey.stackoverflow.co\/2025\/ai\" target=\"_blank\" rel=\"noopener nofollow\">66% of them run into &#8220;AI solutions that are almost right, but not quite&#8221; \u2014 and 45.2% find debugging AI-generated code more time-consuming than writing it<\/a>. That gap between almost-right and right is where engineering work has relocated. Which means the thing you are hiring for has changed shape, and the entry-level candidate you were about to write off is the one most people are mispricing.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Key takeaways<\/h2>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Production of code is no longer the bottleneck. Validation is \u2014 and 30% of developers report little to no trust in AI-generated code.<\/li>\n\n\n\n<li>The cost is measurable in the codebase: duplicated blocks of five or more lines rose eightfold in 2024, while refactoring fell.<\/li>\n\n\n\n<li>Early-career workers in the most AI-exposed occupations have seen a 16% relative employment decline. That is a market fact, not a verdict on capability.<\/li>\n\n\n\n<li>Judgment is observable in weeks, not years \u2014 if you test it with a real task in a real repository instead of a puzzle interview.<\/li>\n<\/ul>\n\n\n\n<h2 class=\"wp-block-heading\">The bottleneck moved, and the metrics show where<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Google&#8217;s 2025 DORA research is unambiguous about the trade. <a href=\"https:\/\/dora.dev\/insights\/balancing-ai-tensions\/\" target=\"_blank\" rel=\"noopener nofollow\">90% of technology professionals now use AI at work and over 80% believe it has increased their productivity \u2014 while 30% report little to no trust in the code it generates, and higher adoption correlates with both higher delivery throughput and higher delivery instability<\/a>. One engineer in the research put it plainly: &#8220;I feel somewhat more productive, but it&#8217;s at a cost. While I end up spending less time writing code, I spend more time babysitting the AI and reviewing what it is trying to do.&#8221;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The residue shows up in the repository. GitClear&#8217;s analysis of <a href=\"https:\/\/www.devclass.com\/ai-ml\/2025\/02\/20\/ai-is-eroding-code-quality-states-new-in-depth-report\/1626250\" target=\"_blank\" rel=\"noopener nofollow\">211 million changed lines of code found that blocks with five or more duplicated lines increased eightfold during 2024, that 2024 was the first year on record where copy\/pasted lines exceeded moved lines, and that moved lines \u2014 the signature of refactoring \u2014 fell 39.9%<\/a>. Assistants make inserting a new block trivial and reusing an existing one unlikely. Nobody on your team wakes up wanting more duplication; it accumulates because writing got cheap and reviewing did not.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">So the scarce resource on a fifteen-person team is no longer keystrokes. It is attention that can look at a plausible diff and say <em>this is wrong, and here is why<\/em>. That is a judgment role, and it is the one you should be recruiting into.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">The &#8220;freshers are disqualified&#8221; argument, and where it breaks<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">There is a confident version of this thesis circulating that ends somewhere darker: if judgment is what&#8217;s being bought, and judgment comes from production scars, then entry-level candidates have nothing to sell. The labour market is behaving as though that is true. Brynjolfsson, Chandar and Chen, using administrative payroll data from the largest US payroll provider, found <a href=\"https:\/\/digitaleconomy.stanford.edu\/publication\/canaries-in-the-coal-mine-six-facts-about-the-recent-employment-effects-of-artificial-intelligence\/\" target=\"_blank\" rel=\"noopener nofollow\">a 16% relative decline in employment for early-career workers aged 22\u201325 in the most AI-exposed occupations, with declines concentrated where AI automates rather than augments<\/a>, while employment for more experienced workers in the same occupations held up.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">That is real, and it is worth taking seriously. But notice what it measures: hiring behaviour under uncertainty, in occupations where the work was automatable. It does not establish that a second-year engineering student cannot exercise judgment. It establishes that employers have no cheap way to tell which ones can \u2014 so they stop buying the category. A market that stops pricing something correctly is a market with an opening in it.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The logical flaw in the harder version of the argument is treating scars as the only source of judgment. Scars are a delivery mechanism \u2014 consequence attached to a decision, fast enough to learn from. There is nothing sacred about acquiring them over six years at a large company. They can be manufactured deliberately, in compressed form, if someone gives a student real work with real consequences and then documents what happened when it went wrong.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"576\" src=\"https:\/\/tierones.io\/blog\/wp-content\/uploads\/2026\/07\/review-gate-judgment-filter-opt-1024x576.jpg\" alt=\"review-gate-judgment-filter\" class=\"wp-image-184\" srcset=\"https:\/\/tierones.io\/blog\/wp-content\/uploads\/2026\/07\/review-gate-judgment-filter-opt-1024x576.jpg 1024w, https:\/\/tierones.io\/blog\/wp-content\/uploads\/2026\/07\/review-gate-judgment-filter-opt-300x169.jpg 300w, https:\/\/tierones.io\/blog\/wp-content\/uploads\/2026\/07\/review-gate-judgment-filter-opt-768x432.jpg 768w, https:\/\/tierones.io\/blog\/wp-content\/uploads\/2026\/07\/review-gate-judgment-filter-opt-1536x864.jpg 1536w, https:\/\/tierones.io\/blog\/wp-content\/uploads\/2026\/07\/review-gate-judgment-filter-opt.jpg 1600w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<h2 class=\"wp-block-heading\">What judgment looks like when you can actually see it<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Abstracted, judgment sounds unhirable. Operationalised, it is a short list of behaviours you can watch for in ordinary work: noticing that output is wrong before being able to articulate why; choosing the boring, legible option when the clever one is unsupportable; asking one clarifying question before writing 400 lines against a misread ticket; escalating at hour two instead of silently guessing until Friday; deleting their own code when the requirement changed. None of these requires a decade. All of them are visible in a pull request and its comment thread.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">What is <em>not<\/em> visible is anything a candidate asserts about themselves. In 2026 the entire self-reported proof stack \u2014 the polished README with a tradeoffs section, the architecture diagram, the confident cost analysis, the metric in the headline \u2014 is generatable in an afternoon. If your process asks a founder to take those artifacts on faith, it is measuring prompt quality. The related failure in interviews, and how to redesign around it rather than police it, is covered in our piece on <a href=\"https:\/\/tierones.io\/blog\/evaluating-ai-assisted-code-interviews\/\">evaluating AI-assisted code in interviews<\/a>.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">A six-week protocol for testing judgment<\/h2>\n\n\n\n<ol class=\"wp-block-list\">\n<li><strong>Pick a task you genuinely want done.<\/strong> Real ticket, real repository, real users downstream. Synthetic exercises test compliance; consequence tests judgment.<\/li>\n\n\n\n<li><strong>Under-specify it slightly, on purpose.<\/strong> Leave one ambiguity in the ticket. Whether they ask about it, and how, is the highest-signal moment in the whole trial.<\/li>\n\n\n\n<li><strong>Say AI tools are allowed, and that you will ask about every line.<\/strong> This removes the incentive to hide usage and reinstates the only thing you care about \u2014 whether they can defend the output.<\/li>\n\n\n\n<li><strong>Review the pull request exactly as you would a teammate&#8217;s.<\/strong> Do not soften it. How someone absorbs a blunt review comment tells you more than the diff does.<\/li>\n\n\n\n<li><strong>Watch time-to-escalation, not time-to-first-commit.<\/strong> The failure mode that costs you money is silent guessing. Log how long they sat on a blocker before raising it.<\/li>\n\n\n\n<li><strong>Run a two-minute walkthrough at the end.<\/strong> Problem, approach, what they rejected and why, what they would do with another week. Cheap to administer, very hard to fake live.<\/li>\n\n\n\n<li><strong>Introduce one real change of direction mid-trial.<\/strong> Requirements move in every startup. Watching someone delete their own work without sulking is a hiring signal in itself.<\/li>\n\n\n\n<li><strong>Write the scars down.<\/strong> One paragraph on what broke, what they did, what they would change. That artifact is the compressed experience \u2014 and it travels with them.<\/li>\n\n\n\n<li><strong>Decide at week six, and mean it.<\/strong> A fixed end date is what makes both sides honest. The structure is laid out in full in our <a href=\"https:\/\/tierones.io\/blog\/six-week-internship-work-trial\/\">six-week work trial guide<\/a>.<\/li>\n<\/ol>\n\n\n\n<h2 class=\"wp-block-heading\">Why this is worth doing rather than waiting<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Because trials convert. NACE&#8217;s 2026 Internship &amp; Co-op Survey found <a href=\"https:\/\/www.naceweb.org\/talent-acquisition\/internships\/intern-conversion-rate-hits-highest-mark-in-five-years\" target=\"_blank\" rel=\"noopener nofollow\">the average intern conversion rate hit 63.1% for 2024\u201325 interns, the highest in five years, with an 88.3% acceptance rate<\/a>. Six weeks of scoped work is not a favour to a student; it is the cheapest evaluation instrument available to a small team, and it produces a hire you have already watched work.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The remaining problem is the one that makes founders skip this entirely: sourcing. That is what Tierones does \u2014 an invite-only network of verified second-year developers from India&#8217;s IITs, NITs and IIITs, where <a href=\"https:\/\/tierones.io\/tier-rank\">Tier Rank<\/a> is built from merged commits, code reviews and shipped projects rather than self-reported claims, and every profile is tied to a confirmed institute email. Verification instead of faith, which is precisely the thing the artifact stack stopped providing. Free while the network grows.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">FAQ<\/h2>\n\n\n\n<h3 class=\"wp-block-heading\">If AI writes junior-level code, why hire a junior at all?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Because the constraint has moved from producing code to validating it. Stack Overflow&#8217;s 2025 survey found 66% of developers cite &#8220;AI solutions that are almost right, but not quite&#8221; as their top frustration, and 45.2% say debugging AI-generated code takes more time. Someone has to close that gap, and senior review capacity is the scarcest thing on a small team.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Doesn&#8217;t judgment require years of production experience?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Some of it does, but not the part you can hire on. Judgment as a working skill means noticing that output is wrong before you can explain why, choosing the boring option under pressure, and escalating early instead of silently guessing. Those behaviours are observable in a few weeks of real work with real consequences. What experience buys is breadth of pattern \u2014 which your codebase and your review comments supply.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">What is the fastest way to test judgment rather than recall?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Give the candidate a real, scoped task in your actual repository and review the pull request as you would any other. Then ask them to walk through one decision they made and one they rejected, in about two minutes. Puzzle interviews test recall, which AI has commoditised; a pull request plus a walkthrough tests the reasoning behind the code, which it has not.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><em>Hiring remote engineering help this quarter? <a href=\"https:\/\/tierones.io\/employers#hire\">Tell us what you need<\/a> \u2014 we&#8217;ll show you the right 40, not the loudest 4,000.<\/em><\/p>\n","protected":false},"excerpt":{"rendered":"<p>66% of developers say their top AI frustration is code that&#8217;s almost right. Judgment is the job now \u2014 and it is testable in six weeks, not six years.<\/p>\n","protected":false},"author":1,"featured_media":183,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[3],"tags":[120,104,123,122,119,57,121,103],"class_list":["post-177","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-for-employers","tag-ai-assisted-development","tag-code-review","tag-engineering-leadership","tag-entry-level-hiring","tag-hiring-junior-developers","tag-remote-interns","tag-technical-debt","tag-work-trial"],"_links":{"self":[{"href":"https:\/\/tierones.io\/blog\/wp-json\/wp\/v2\/posts\/177","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/tierones.io\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/tierones.io\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/tierones.io\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/tierones.io\/blog\/wp-json\/wp\/v2\/comments?post=177"}],"version-history":[{"count":1,"href":"https:\/\/tierones.io\/blog\/wp-json\/wp\/v2\/posts\/177\/revisions"}],"predecessor-version":[{"id":185,"href":"https:\/\/tierones.io\/blog\/wp-json\/wp\/v2\/posts\/177\/revisions\/185"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/tierones.io\/blog\/wp-json\/wp\/v2\/media\/183"}],"wp:attachment":[{"href":"https:\/\/tierones.io\/blog\/wp-json\/wp\/v2\/media?parent=177"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/tierones.io\/blog\/wp-json\/wp\/v2\/categories?post=177"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/tierones.io\/blog\/wp-json\/wp\/v2\/tags?post=177"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}