The Judging Playbook — What Actually Happens to Your Submission
Derived from: the 2024-25 Innovation Stage Judge Guide, the Scoring Guide, the Chief Judge Q&A (Simon Glinsky), and close reading of five real Innovation Briefs assigned for judging in 2025 (Greencrete, SUNSYNC, Ocean Energy Dynamics, P-Bump, Puppy WC — all Energy & Environment, Global chapter).
This document exists to answer one question: what separates a 2 from a 4?
Part I — The judge's actual situation
Internalise this before writing anything.
Your judge is a volunteer professional — an entrepreneur, VC, engineer, scientist, or educator. They have 5 teams and 5–10 total hours. That is ~90 minutes for you, including writing five prose comment fields. They are explicitly instructed to:
- Google your claims. "Perform an online search to verify originality of the approach or innovation."
- Read the whole submission before scoring, and to re-read across multiple sessions.
- Watch the video and open the website.
- Open your references PDF.
- Score each theme independently — a broken innovation can still earn 4 on Marketing.
- Write as a coach, not a grader.
- Never contact you. Every question they have goes unanswered — and becomes a doubt.
The five judge questions, in the order they get asked
- (Innovation) Have I seen this before? → they search
- (Practicality) Would this actually work? → they apply physics/domain sense
- (Storytelling) Do I believe these people? → they read for polish and coherence
- (Marketing) Do they know who buys this? → they look for named customers and named competitors
- (Finances) Do the numbers add up? → they check arithmetic, leniently
The single most important consequence: because the judge cannot ask you anything, every ambiguity resolves against you. "I am left wondering" is the phrase that costs points.
Part II — The five briefs, scored and diagnosed
I read these as an assigned judge. Names as submitted. Scores below are my judgments applying the official rubric, not official Conrad scores.
Calibration table
| Brief | Innov (30) | Story (20) | Pract (20) | Mkt (20) | Fin (10) | ~Total | Verdict |
|---|---|---|---|---|---|---|---|
| Ocean Energy Dynamics | 3 | 5 | 2 | 5 | 4 | ~76 | Superb business writing, physics problem |
| P-Bump | 3 | 4 | 2 | 4 | 3 | ~66 | Real engineering, fatal energy-accounting error |
| SUNSYNC | 2 | 3 | 4 | 3 | 3 | ~58 | Real prototype, unoriginal innovation |
| Puppy WC | 1 | 3 | 2 | 3 | 2 | ~42 | Fluent prose, product already exists |
| Greencrete | 1 | 1 | 1 | 1 | 1 | ~20 | Unfinished |
None of these is a finalist. Finalists need ~85–90+. Note what does not correlate with the total: word count, prose fluency, and enthusiasm.
Case 1 — Greencrete (~20/100): the anatomy of a 1
Low-carbon concrete using biomimicry. Two students.
What a judge sees immediately:
- Q2 (Team) describes skiing and basketball hobbies and never states a role. 150 words spent on nothing.
- Q7 (Competition) answers the wrong question — it repeats target-customer text instead of naming competitors.
- Q8 (Go-to-Market) is copy-pasted verbatim from Q6.
- Q9 says the business model is "listening" — a typo for licensing, uncorrected.
- Q5 (Validation) opens: "we don't really have a complete 100% solution."
- Q10 (Fundraising): total development cost "5–10 thousand dollars" to commercialise a new cement chemistry.
- Innovation image: a Tinkercad tree on a brown box.
- Website: a raw Canva /edit link — a judge clicking it may land in an editor, not a site.
Diagnosis: this is not a bad idea, it is an unfinished submission. Low-carbon concrete is a legitimate multi-billion-dollar problem; Holcim is correctly named. The team had a real thread (biomimicry → spiral fibre reinforcement) and abandoned it.
What the Chief Judge says to do here: score it honestly — 1s and 2s across the board — but comment as a coach. "It's reasonable for this team to receive one or two-point scores... That fulfills goal #2. Our #1 goal is to create learning for the students, including those who underperform."
Teaching value: highest of the five. Every failure is mechanical and fixable in an afternoon: answer the question asked, don't duplicate answers, proofread, publish the website properly.
Case 2 — Puppy WC (~42/100): fluent writing cannot rescue a non-innovation
A self-cleaning, flushable pet toilet with sensors. Two students, South Korea.
The prose is genuinely the second-best of the five: clean paragraphs, a systematic competitive comparison against puppy pads, washable pads, and artificial grass mats.
And it scores 1 on Innovation, because the judge is required to run an originality search — and automatic self-cleaning flushing pet toilets are an existing consumer product category (Inubox, BrilliantPad, and others). The brief never mentions that these exist. It positions only against pads and mats, which is the comparison set that makes the product look novel.
Compounding it:
- No engineering specificity anywhere — no sensor type, no water volume per flush, no power draw, no cost.
- "Proprietary flushing mechanism... engineered to accommodate the specific consistency of pet waste" — asserted, never described.
- Quantification is a range with no basis: "hundreds to over a thousand disposable pads per year."
The lesson, and it is the single most important one in this document:
Choosing a favourable comparison set is the most common way teams destroy their own Innovation score. The judge picks the comparison set, not you. If a competitor exists and you don't name it, the judge concludes you either didn't look or you're hiding it. Both are worse than the competitor.
The fix is not a different product — it is Q7 naming the real incumbents and then earning the differentiation honestly ("existing units cost $500+ and require proprietary cartridges; ours...").
Case 3 — SUNSYNC (~58/100): a working prototype with nothing new in it
Solar panel that tracks the sun (photoresistors + servo + Arduino) and concentrates light with convex lenses. Three students, City of Knowledge Academy, Nigeria.
This team did more real work than any other in the set and scores mid-table. That is the lesson.
Strengths a judge genuinely rewards:
- A built, tested prototype. Practicality = 4.
- Real validation prose: they describe covering one light sensor to verify servo response, and checking the 9 V panel against the Arduino's 5 V/3.3 V logic to avoid damage. That is authentic engineering process, and it reads as authentic.
- Clean structure, cited sources, clear roles, SDG 7 framing.
Why Innovation = 2:
- Solar tracking is 1960s technology; concentrated photovoltaics is a mature field. Both are the first hits of the originality search the judge is required to run.
- The combination is not defended as novel — no argument for why tracker + lens together is more than the sum.
- "We used a unique method of programming the Arduino board so our code can't be replicated." This one sentence is actively damaging. To an engineer-judge it signals a fundamental misunderstanding of software, of reverse engineering, and of what IP is. IP Defensibility is an explicit sub-criterion; this answers it wrongly and confidently.
- "Generates twice as much energy as normal solar panels, up to 90%" — two incompatible claims in one sentence, no measurement, despite the team owning a working rig that could have measured it.
Finances: $120 prototype → $1,200 commercial unit, with no bridge explaining the 10×. "$500 for the initial phase" of rollout is not a rollout budget.
Diagnosis: they had the instrument to answer the question and didn't run the experiment. A single afternoon logging watt-hours from a tracked+lensed panel versus a fixed panel, plotted, would have moved Innovation and Practicality both — and turned a marketing claim into evidence.
The SUNSYNC rule: if you built it, measure it. An unmeasured prototype is worth less than a well-reasoned design, because it proves you had access and chose not to look.
Case 4 — P-Bump (~66/100): rigorous arithmetic, wrong physics
Piezoelectric speed bump: hydraulic 10:1 force amplification onto a PZT stack, for Indonesia. Three students.
The most technically ambitious brief of the five. Genuine specifications: 10 cm → 31.6 cm cylinder bore (correctly giving 10:1 area ratio), a 4-section PZT stack of 100 layers each, 28×28 cm base, vacuum return mechanism, inline URLs to ScienceDirect and traffic-authority data, a named real competitor (the Shibuya Station piezoelectric floor) with honest advantages and disadvantages, a utility-patent + trade-secret strategy, and a coherent B2G model with Power Purchase Agreements.
This is what a 4–5 on Marketing and Practicality is supposed to look like structurally.
And Practicality still scores 2, for one reason a judge in this field will catch in 30 seconds:
A speed bump does not harvest free energy. It harvests energy from the vehicle's engine.
Depressing the bump does work on the vehicle — extra rolling resistance, extra fuel burned, at internal-combustion efficiency (~25%) and then piezo conversion. The device is a net energy loss machine that quietly taxes every driver. The brief never mentions this. Every credible piezo-road pilot has failed on exactly this economics.
Then the arithmetic, which looks rigorous and isn't:
- Claimed stress on the PZT stack: 46,875 Pa. That is ~0.05 MPa. PZT is typically driven at tens of MPa. The stack is loaded at roughly a thousandth of a useful stress — the geometry defeats the amplification.
- Energy per car computed as
E = P·Vwith a 1 mm displacement → 4,113 J per car, then scaled by 20,000 cars/day to "power a household." The 1 mm displacement of a 28×28 cm ceramic stack is not physically available at that stress, and the formula conflates pressure-volume work with recoverable electrical energy at 88% "efficiency" of the material.
The lesson — and it is the opposite of the Puppy WC lesson:
Numbers do not create credibility. Numbers create checkable claims. A specific wrong number is worse than an honest range, because it is falsifiable and the judge is instructed to falsify it.
The redemption path was available and cheap: a sanity-check paragraph. "Energy is drawn from the vehicle, so this is only viable where the bump is already required for traffic calming and the marginal fuel cost is accepted as the price of the safety feature." That reframing is honest, survives scrutiny, and is more interesting.
Case 5 — Ocean Energy Dynamics (~76/100): the best business writing, the worst physics
Wave-energy system integrated into a ship's hull: one-way inlet, debris mesh, turbulence-reducing grids, lead oscillators on slide rods, sealed hydraulic cylinders, carbon-fibre hydraulic turbines. Four students.
This is what finalist-grade Marketing, Finance and Storytelling writing looks like. Study these:
- Named real competitors: CalWave, Mocean Energy, Carnegie Clean Energy — and a precise positioning gap ("fixed-location systems... entirely unsuitable for moving vessels").
- Segmented customers with a payer/buyer distinction: yacht owners (primary), shipbuilders (secondary), charter companies (tertiary), defence (adjacent) — each with its own reason to buy.
- Sized market with a trajectory: $240 B (2023) → $460 B (2033) shipbuilding.
- Real unit pricing with tiers: $50,000 single-unit for small yachts; 100, 000–250,000 multi-unit.
- Recurring revenue: maintenance contracts, licensing, premium monitoring.
- Use of funds as percentages: $10 M split 20% R&D / 40% manufacturing / 20% pilots / 20% marketing.
- Layered IP strategy: patents on grids and hydraulic loop, trade secrets on manufacturing and material composition, trademarks for brand.
- Honest disadvantage stated: higher initial cost, then answered with segment logic (wealthy, image-conscious buyers).
- Attachments used properly: a 3D model, a rendered animation, and a partial physical prototype (turbine + DC motor + Arduino) with the limitation stated plainly ("we could not use water because the motor was not waterproof").
That last point deserves emphasis: they stated what their prototype could not do. Judges reward this. It is the opposite of SUNSYNC's unmeasured 90% claim.
So why does Practicality score 2 and Innovation only 3?
A ship's own motion in waves is the energy source. Extracting energy from that motion adds drag and increases the ship's propulsion demand. The brief then states the recovered electricity "powers an innovative propulsion system" — which closes the loop into a perpetual motion machine.
A marine-engineering judge sees this instantly. There is a legitimate version of this idea — parasitic harvesting from wave motion for hotel loads on a moored or drifting vessel, where the drag penalty is irrelevant because the ship isn't going anywhere. That version is defensible, still novel-ish, and would have scored 4 on Practicality with the same writing.
Diagnosis: the team out-wrote their engineering review. No adult with a fluid-dynamics background read Q4 before submission. Everything else was finalist-grade.
Part III — The cross-cutting patterns
The four failure archetypes
| Archetype | Symptom | Example | Cost |
|---|---|---|---|
| Unfinished | Duplicated answers, typos, wrong question answered, broken links | Greencrete | Everything |
| Favourable comparison set | Positions only against weak alternatives; real incumbent unnamed | Puppy WC | Innovation → 1 |
| Unmeasured prototype | Built the thing, never quantified it; claims unbacked | SUNSYNC | Innovation, Practicality |
| Polish over physics | Excellent business writing wrapping an unsound mechanism | Ocean Energy, P-Bump | Practicality → 2, caps total ~75 |
The last one is the most dangerous, because it is invisible to the team, to a business-trained coach, and to any AI writing assistant. It is only caught by a domain expert reading for mechanism.
The Innovation score is decided by one paragraph
Innovation is 30% — the largest single block. The Chief Judge's stated test:
"If anyone can copy that application, it's not a strong innovation."
And the worked example: a new soda flavour scores 1/5. The same new flavour, if it reduces tooth decay, scores 3–5 — "a worthy and protectable innovation, even though it's a new application of an existing product."
So Innovation is not measured by technological sophistication. It is measured by defensible differentiation. SUNSYNC used more advanced hardware than the soda example and scored 2.
The sub-criteria say this outright: Verification (did an online search rule out duplicates?) and IP Defensibility (patent, trade secret, copyright, first mover, contracts, ecosystem capture). Note that the last three are business moats, not legal ones — a team with no patentable technology can still score well by owning a distribution relationship or a dataset.
Practicality does not require a prototype — and this is under-exploited
The Chief Judge is unambiguous:
"Teams can attain 5 points for Practicality without a prototype."
The cited example: a jet-engine electrolyser hydrogen burner — impossible to prototype on a student budget — earned 5 on Practicality and 4–5 on Innovation via component-technology explanation, credible drawings, technical explanation, cited expert testimony, and exploratory talks with manufacturers.
The evidence ladder (any one or more establishes proof of concept):
- Existing applications of the component technologies
- Expert testimony — a named person with credentials who reviewed it
- Research verifying feasibility — cited literature
- Convincing graphic representation — CAD, schematic, animation
- Partial or full prototype or demonstration
- Describing further research or experiments that would verify feasibility
Rungs 1–4 and 6 cost no money. Rung 2 costs one polite email. Most teams skip straight to "we couldn't afford a prototype" and leave 20% of the score on the table.
Finances: the cheapest points in the competition
Judges are told to be lenient here and it is only 10%. The Chief Judge's own scale:
| Score | What earns it |
|---|---|
| 2–3 | A reasonable effort; financials that add up to totals and show expected expense categories |
| 4 | Sound internal logic — even if actual market costs weren't surveyed |
| 5 | Strong logic, strong pricing strategy, and unit profit or supporting research/comparison |
"A terrific team and product should not be knocked out of final consideration solely due to financial projections — as long as they made an effort with some logic involved."
Translation: 4/5 on Finances is available to any team that computes one unit's cost, one unit's price, and shows the arithmetic. Ocean Energy got there. Greencrete's "$5–10 thousand" did not.
The classic student error the Chief Judge names: costing only the immediate things (prototype materials) and never the full development and deployment.
Storytelling is where the video and website actually count
Storytelling & Professionalism (20%) explicitly "encompasses questions 1, 2 + video + entire submission and all attachments." Its sub-criteria include Credibility Boost (do the video and website reinforce expertise?) and Polish & Consistency.
This is why Greencrete's Canva /edit link and Tinkercad tree are not cosmetic issues — they are scored. And it is why judges are told to correct spelling in their own comments: the competition takes professional presentation seriously in both directions.
Concretely, judges reacted badly to casual video narration:
"'Yeah, this is like three-stage, this is like seven-stage or something, I don't know, I didn't count it' significantly lowers my belief that you'll champion this product and use investor resources well."
What judges say about exaggerated claims
If a team claims sales or deployments that seem inflated, judges are told not to award points for business achievements at all — the rubric doesn't reward having made a sale — and to raise it dispassionately:
"5,000 units manufactured means a total cost of $2M; you also need $1M for your building complex; the Nike deal is worth $800,000. Do you mean these have already been achieved, and if so, where and how did you fund the $3M required investment? Or are these projections for the future? I am uncertain."
Implication for teams: never inflate. It converts a neutral section into an active credibility problem, and it is trivially detected by arithmetic.
Part IV — The improvement ladder
If you are a team (or a coach) with a draft, work these in order. Ordered by points-per-hour.
Tier 0 — Mechanical (2 hours, worth ~15 points at the low end)
- Every question answers the question asked. Re-read the prompt, then your answer.
- No duplicated text between questions.
- Spellcheck. Then read aloud.
- Website is published and opens in an incognito window. Video link plays without login.
- References PDF exists, is complete, and includes any AI tools used.
Tier 1 — The originality search (3 hours, decides your 30% block) 6. Search for your innovation as a product category, not as your specific design. Search in the language of the industry, not the language of students. 7. Find the three strongest existing solutions — including commercial products, not just papers. 8. Name them in Q7. State honestly what they do better. 9. Rewrite Q4 to defend the delta — what specifically is new, and why can't the incumbent just do it? 10. Answer IP Defensibility with a real mechanism: patent claim, trade secret, dataset, first-mover contract, ecosystem lock-in. Not "we'll copyright our code."
Tier 2 — Mechanism review (1 expert-hour, prevents the ~75 ceiling) 11. Find one adult with domain expertise — not your coach, not a business person — and ask them one question: "Is there a reason this can't work?" 12. Specifically audit energy and mass balance. Where does the energy come from? Is anything a closed loop? Does conservation hold? 13. If there's a flaw, reframe rather than abandon. The honest constrained version usually scores higher than the ambitious impossible one.
Tier 3 — Evidence (5 hours, worth up to 20%) 14. Climb the evidence ladder as far as budget allows. Aim for ≥3 rungs. 15. If you have a prototype, measure something and plot it. 16. Email 2–3 real experts or potential customers. Quote them with names and titles in Q5. 17. State your prototype's limitations explicitly.
Tier 4 — Business arithmetic (4 hours, worth ~10–20 points) 18. One unit: bill of materials → unit cost → price → unit margin. Show the table. 19. Full development cost to market, not just prototype cost. 20. Market size with a source and a date. 21. Segment customers and identify where buyer ≠ payer. 22. Use of funds as a percentage split.
Tier 5 — Narrative (3 hours, worth up to 20%) 23. Q1 Elevator Pitch: problem → mechanism → who buys → why now. 150 words, no adjectives. 24. Q2 Team: roles and relevant capability only. Delete hobbies. 25. Video: rehearse. No "I don't know." Show the model doing something. 26. Website: consistent brand, working links, the same model image as the video.
Part V — Self-scoring rubric card
Score yourself 1–5 per theme, honestly. Multiply: Innovation×6 + Storytelling×4 + Practicality×4 + Marketing×4 + Finances×2 = /100.
Innovation — Can I name the 3 strongest existing solutions and state my defensible delta? Would a stranger's 10-minute Google search find something that makes me look uninformed? Do I have a real moat?
Storytelling — Would an investor read past Q1? Are the video, website, and brief telling one consistent story? Is there a single typo, broken link, or unpublished page?
Practicality — How many rungs of the evidence ladder do I occupy? Has a domain expert confirmed no mechanism-level flaw? Does energy/mass balance hold?
Marketing — Have I named real competitors and real customer segments? Do I know who pays versus who uses? Is my market size sourced?
Finances — Can I state the cost and price of exactly one unit, and the margin? Does my development budget cover the full path to market? Do the numbers add up?
Below 60: finish the submission (Tier 0–1). 60–75: you likely have a mechanism or originality problem (Tier 2). 75–85: you have a good project; evidence and unit economics are the gap (Tier 3–4). 85+: finalist territory. Now rehearse the pitch.