The Discipline of the Empty Cell: Why Cricket Analytics Must Learn to Say 'Insufficient Information'
**মূল উত্তর:** ক্রিকেট বিশ্লেষণে "তথ্য অপর্যাপ্ত" একটি বৈধ ও সিদ্ধান্ত-উপযোগী ফলাফল। ন্যূনতম নমুনা ও স্পষ্ট সংজ্ঞা ছাড়া কোনো খেলোয়াড়ের সংখ্যা সিদ্ধান্তের ভিত্তি হতে পারে না; অজানা তথ্য মানে খারাপ পারফরম্যান্স নয়। **মূল তথ্য:** - ২০১৭ সালে ময়মনসিংহে বিপিএলের ১,২৪০টি শট হাতে ট্যাগ করে এক্সজি মডেল তৈরি হয়েছিল। - ২০২০ সালে ১৮ ম্যাচে খালি Stadiumে হোম এক্সজি ০.৩৪ কমে ও পিপিডিএ ২.১ বাড়ে। - ২০২২ কাতার বিশ্বকাপে মরক্কোর লো-ব্লক প্রতি শটে ০.৫৪ এক্সজি ছাড়ে; হাকিমি ১১.৮ কিলোমিটার দৌড়ান। - ২০২৫ সালে ৩৩ বছরের এক মিডফিল্ডারের ইনজুরি ঝুঁকি ৩৮ শতাংশ ধরা হয়; মিনিট কাটার পর পেশির ইনজুরি ৪০ শতাংশ কমে। - ৪২ বলের নমুনায় ব্যাটসম্যানের স্ট্রাইক রেট সিদ্ধান্তের জন্য যথেষ্ট নয়। **সূত্র:** Stage-2 ডিপ প্রফেশনাল অ্যানালাইসিস কাঠামো নথি (তারিখ নথিতে অনুল্লেখিত) | Cross-checked: cricsultan.com **সম্ভাব্য Next প্রশ্ন:** প্রশ্ন: কত বলের নমুনা হলে একজন বোলারের Economy সিদ্ধান্তে ব্যবহার করা যায়? উত্তর: সাধারণত কয়েকশো বল দরকার, এবং উইকেটের ধরন ও ম্যাচ-পরিস্থিতি অনুযায়ী আলাদা করতে হয়। প্রশ্ন: ট্রান্সফার উইন্ডোতে কোনো খেলোয়াড় সম্পর্কে তথ্য না থাকলে ক্লাব কী করবে? উত্তর: ন্যূনতম-নমুনা নিয়ম লিখে রাখা, ফাঁকা ঘরের খাতা প্রকাশ করা এবং ট্রায়াল ক্যাম্পের মতো যাচাইযোগ্য পদক্ষেপ নেওয়া। প্রশ্ন: ক্রিকেট ডেটার নির্ভরযোগ্যতা যাচাইয়ের সবচেয়ে কার্যকর উপায় কী? উত্তর: প্রতিটি এন্ট্রির সংজ্ঞা, সময় ও সূত্র লিপিবদ্ধ রাখা, যাতে cricsultan.com-এর মতো প্ল্যাটFormে ডেটা পুনরায় যাচাই করা যায়।
The call came at half past eleven at night. On the other end, a franchise's cricket-operations man. Three words: "Do we sign him?" On my screen sat a table. A nineteen-year-old left-arm spinner, four matches, two of them rain-shortened, one with no ball-by-ball data uploaded at all. The column that mattered most — his line-and-length consistency in the powerplay, his economy against left-handed batters — was blank. I stayed quiet for seven minutes. What I said next was not a recommendation: "If you can answer three questions for me, I can tell you. Right now I don't know."

Empty cells are not new to me. In 2026, hand-tagging 1,240 Bangladesh Premier League shots from a desk in Mymensingh, I saw blank columns every day. The hardest part of cricket analysis is not building statistics; it is deciding when not to produce a number.
Where the gaps are born
Modern cricket analysis runs in two stages. The first extracts information points from raw events: who faced how many balls, who bowled which over, what happened on which delivery. The second arranges those points into a structure that means something. The problem is that when the first stage comes back empty, the pressure lands on whoever is sitting at the second stage to fill the gap. The market has no price for a blank answer. Agents, broadcasters, fantasy platforms, social media — everyone wants a number. Nobody asks where the number came from.
We are in the middle of a transfer and squad-building window. This is when the pressure to fill blank cells peaks. Every transfer rumour is a data point with a heartbeat. But a heartbeat is not the same thing as informational reliability. The structure of the release clause, the space it takes in the wage bill, the agent's commission architecture — that is the real story, and that is exactly where the darkest room sits.
Twice in my own career I have had to fight empty data, for two very different reasons. In 2026, at a World Cup desk in Russia, I was mapping Croatia's pressing. Every match was codable, because every match in that tournament carried tracking data of the same standard. Croatia's PPDA of 8.7, Luka Modric's 13.1 kilometres in one match — those numbers only meant something because the definition was identical across every game.
Then in 2026, working with Sheikh Russel during the COVID hiatus, the picture inverted. Across 18 matches in empty stadiums, home xG fell by 0.34 and PPDA rose by 2.1. Empty stadiums taught me that home advantage is a social contract, not a table line. We used that 18-match sample to set up a low-block 5-3-2; over the final five matches we conceded 0.8 xG per game and avoided relegation. But even there, one section of my first report was blank: how far a player's nerve collapses in an empty ground is not something any xG model measures.
A mechanism map of missing data
From years of watching matches and building tables, I have come to believe blank cells arrive from five specific places. Each needs different treatment.
First, uneven coverage. Large parts of domestic and age-group cricket have no ball-by-ball data. When a player gets a national call-up, the raw material of his previous three seasons is stored nowhere. The sample sitting in a selector's hands is therefore not a full sample — it is a selection truncation, and it is larger than the data itself.
Second, format-label errors. A player's T20 economy and ODI economy get averaged together, even though the two games have different over budgets, field settings and risk profiles. Without splitting formats, the analysis is close to useless.
Third, injury and workload opacity. How many balls a bowler has delivered can be counted; how much his shoulder hurts and how well he slept cannot.
Fourth, missing environmental variables. Without coding pitch character, dew, travel time and match breaks, a home-away gap gets misread as a talent gap.
Fifth, inconsistent definitions. When two scouts use the phrase "line and length" to mean two different things, their outputs cannot be summed.
The real cause of a blank cell is usually not a shortage of information; it is a shortage of discipline in definition, coverage and disclosure.
'Unknown' and 'false' are not the same thing
Here is my central argument. We may know nothing at all about a player, and that does not mean he is bad. Unknown means unknown. This ordinary distinction gets erased in cricket conversation almost every time. When someone says "we have no data on him", the listener hears "he is bad".
An example. Say a left-handed batter has 42 balls of data against a leg-spinner. If his strike rate across those 42 balls is 140, that decides nothing. A batter's strike rate against a spinner generally needs several hundred balls to stabilise, and even then it must be split by pitch type and match situation. Signing or dropping a player on 42 balls carries the same risk in both directions.
An analysis that does not show its own uncertainty is not analysis; it is a bet, dressed up in a tidy table. That is why my reports always carry a confidence interval — and always carry the sentence stating what evidence would change my mind. That second sentence does the most work, because it lets a colleague interrogate the reasoning rather than the conclusion.
What 64 coded matches taught me
At the 2026 World Cup in Qatar I coded all 64 matches for a South Asian scouting network — PPDA, xG, progressive passes. Before Morocco versus Spain, the model showed Morocco's 5-4-1 low block conceding only 0.54 xG per shot, with Achraf Hakimi covering 11.8 kilometres to hold the rest-defence. Morocco won on penalties. My model had not predicted the win. The model did not predict this; it only made the surprise legible.
Morocco did not break the model; they exposed the variables we had been too lazy to name — compactness, pressing triggers, the management of space after restarts. That is the analyst's actual job: not to turn surprise into prophecy, but to make surprise intelligible.
And where data really is strong, it saves careers. In 2026 I told an Asian club that a 33-year-old midfielder carried roughly a 38 percent muscle-injury risk if he kept playing that calendar. The club cut his minutes. Muscle injuries fell 40 percent and the side reached the knockout round. That is the best version of the loop: measurement to decision, decision to outcome. I delivered that report two days late, though, because I was re-checking every model input. That perfectionist habit is my weakness, and I now schedule time around it.
The reverse trap: using uncertainty as a shield
There is a danger here, and it works against my own instincts. "Insufficient information" can be an honest answer, or it can be a place to hide. There are two distinct failures.
The first: analysis that sells certainty, stuffing the blank cell with vibes. The second: the analyst who never says anything at all, dodging decisions under the banner of uncertainty. Both cheat the reader, because the reader is looking for help, not applause.
What separates them? Declaring your limits, and still leaving the door to a decision open. If I say "four matches tell us nothing", I must add: "but here are three things we can verify this week — a trial camp, five overs in a second-team match, a fitness test." Uncertainty converted into a question is not paralysis; it is work.
This is where the integrity of the data chain matters. The promise blockchain makes in financial sectors — that you can walk backwards and see who entered what, when, and under which definition — is precisely the auditability cricket's data pipelines need. I am not claiming the technology is the solution. I am claiming that if the links in the data chain are not transparent, even the best model will only produce well-arranged errors. Every piece of my writing carries footnotes and a methodology note, because the reader has a right to audit the claim.
What to do in this window
For clubs and boards in this squad-building window, my advice is plain, and there are three parts. One, write down a minimum-sample rule — below a certain number of balls or innings, no number appears beside a player's name. Two, publish a "blank-cell ledger" alongside scouting reports, stating what we do not know and why. Three, keep at least one person in the room whose only job is to say: "We do not have enough evidence for this decision."
The blog in Mymensingh was my first stadium: no crowd, only signal. That signal still says the same thing — the power of data lies not in its volume but in its discipline.
So the question is yours: for the player you are signing in this window, which evidence can you actually produce — and if you cannot produce it, how will you make the decision?
