International FootballThe Empty Spreadsheet and the Trap of a Clean Report
International Football

The Empty Spreadsheet and the Trap of a Clean Report

core_answer: Phân tích bóng đá thất bại nghiêm trọng nhất khi dữ liệu trống bị đọc như sự an toàn. Một ô trống là trạng thái chưa được đánh giá, hoàn toàn khác trạng thái không có rủi ro. Nhà phân tích phải đếm ô trống trước khi đọc số và tuyên bố rõ giới hạn của mình.
key_facts: Croatia tại World Cup 2018 chỉ để đối phương chạm bóng trong vòng cấm 4,2 lần mỗi trận, dù giữ bóng trung bình 58%.; Năm 2017, một câu lạc bộ hạng Nhì Thượng Hải phát hiện hậu vệ phải đối thủ dâng cao 12 mét, để lộ khoảng không 25 mét.; Tỷ lệ kiểm soát bóng 63% với 612 đường chuyền có thể chỉ tạo ra 9 đường chuyền vào vòng cấm trong 90 phút.; Sân trống năm 2020 loại bỏ áp lực khán đài và làm lộ rõ khác biệt giữa chiến thuật bài bản và chiến thuật cảm hứng.; Trạng thái trống trong báo cáo thể lực từng khiến một câu lạc bộ ký hợp đồng với cầu thủ có vấn đề đầu gối chưa được chuyển hồ sơ.
source_attribution: Phân tích chuyên môn của Andrew White, Thượng Hải | Cross-checked: VuaBong.vn
related_qa: q: Vì sao một báo cáo dữ liệu sạch vẫn có thể là dấu hiệu xấu?, a: Vì báo cáo sạch có thể phản ánh một quy trình thu thập dữ liệu đã ngừng hoạt động, chứ không phải một hệ thống thật sự lành mạnh.; q: Tỷ lệ kiểm soát bóng có đáng tin không?, a: Không, vì con số này gộp chung hai thứ khác nhau là kiểm soát vị trí bóng và kiểm soát nhịp độ trận đấu.; q: Bộ lọc nào giúp đánh giá một tin đồn chuyển nhượng?, a: Bốn câu hỏi về nguồn tin, vị trí thay thế, cấu trúc hợp đồng và trách nhiệm nếu thương vụ đổ vỡ.

2:40 a.m. in Shanghai. Eleven monitors still glow in the editing room, the cooling fans hum like a match played in an empty stadium, and on my spreadsheet there is one empty column. That column was supposed to hold the number of touches inside the penalty area for a team I was tracking. It was empty, not because the team never touched the ball there, but because the footage I needed had not arrived that night. I sat there, hand still on the mouse, and recognised the choice every analyst faces at least once: keep writing from imagination, or stop and say there is nothing to say yet. I chose the second. It took me several more years to understand why the second choice is so hard, and why it is the real test that separates an analyst from a salesman. In 2026 I worked as a technical consultant for a second-tier club in Shanghai. Our budget was roughly one fifth of our rivals'. No analysis department, no wide-angle camera system, no motion-tracking software. A battered computer, a hard drive of footage borrowed from a local broadcaster, and six weeks. I spent those six weeks cutting every match of our opponent, measuring the distance between their lines with a ruler on the screen, logging every time their right-back pushed up. By week five the number had stabilised at a suspiciously consistent 12 metres. Whenever they attacked, the right-back advanced 12 metres from his defensive base, and the space behind him was 25 metres wide. I circled that patch of ground in red on the whiteboard and proposed a 4-3-3 variant with the left winger running diagonally into that corridor. The match finished 3-0, all three goals originating from the red circle. That story is usually told as a victory for data. What I remember most is the night before the match, when the footage of the opponent's two most recent games failed to download. I had four matches, was missing two, and I knew that if the opponent had changed how they operated in those two, my entire conclusion would collapse. I wrote a short note to the coaching staff stating clearly that the model rested on four samples, that two were missing, and that the probability of my conclusion being wrong sat somewhere between low and moderate. Nobody on the coaching staff read that note carefully. They read the conclusion. That was my first lesson about data gaps. People do not read the part where you admit you do not know. They read the part where you assert. Years later, in the technical operations room of a Shanghai broadcaster during the 2026 World Cup, I met the same problem at a larger scale. After the group stage I rewatched all fourteen Croatia matches. I counted how often opponents touched the ball inside their penalty area: 4.2 times per match. That number is absurdly low for a team that reached the final. The familiar explanation is that Croatia defended in numbers, parking a bus in front of goal. The footage did not show that. Croatia averaged 58 percent possession and deliberately slowed the tempo in the final ten minutes of each half, not to protect a scoreline but to break their opponent's rhythm, forcing them to restart from scratch every time they won the ball back. I called it tactical breathing and wrote a three-thousand-word piece. The editor said it was too academic. I redrew the whole thing as a diamond diagram with arrows and turned it into a story about a machine built to torment opponents. That piece was read. What I learned was not that data needs pictures. What I learned is that a number only lives when it is attached to a specific decision by a specific person on the bench. Now back to that empty column. An empty cell in a data report is more dangerous than a wrong number. A wrong number can be caught, argued over, corrected. An empty cell cannot. It drifts quietly through the system, and by the time it reaches the final reader it has become a blank that the human mind automatically fills with the most convenient assumption. Hand a coach a data sheet on an opponent with three blank cells and he will not think he is missing data. He will think the opponent has no weaknesses in those three areas. This is the simplest and most damaging psychological mechanism in the trade: people read silence as safety. Medicine has a name for this error — a false negative is treated as good news until the patient returns with a late-stage tumour. Football's equivalent: no data on the opponent's left-sided defending, so we default to the left side being fine. Across years of writing and consulting I have sorted silence into four kinds, each demanding a completely different response. The first is silence from missing collection. This is the most common and least serious. Footage did not arrive, cameras did not cover enough, the league publishes no detailed data. The response is to state plainly that the data does not exist and never to infer from the emptiness. The second response is a blunt question: if I did not have this data, which of my conclusions would change? If the answer is none, you do not actually need that data and you are wasting time. If the answer is all of them, you are not holding an analysis. You are holding a hypothesis. The second kind is silence from hidden data. In many leagues, off-ball running is not published, pressing actions are not counted correctly, distances between lines are measured by nobody. This silence is not because the data does not exist but because nobody pays to measure it. With this kind the only fix is to measure it yourself. My six weeks of cutting tape in 2026 was the answer to this kind of silence. No software, no data vendor, just a ruler, eyes and time. I still hold that an analyst willing to spend forty hours measuring something nobody wants measured has a greater edge than an analyst with access to every database on earth. The third kind is silence from late data. In modern football this is especially severe during transfer windows and in congested calendars. You have data from three rounds ago, but the club changed coach last week. You have season-long data, but the first-choice centre-back is injured. Old data inside a new system is wrong data, and it is dangerous because it looks right. The fourth kind, and the one I want to dwell on, is silence from misread data. This is the possession case. I have said this at many seminars and been argued with every time: possession percentage is the most deceptive of all the metrics television puts on screen. A team farming 60 percent of the ball through sideways passes inside its own half is not controlling the match. It is holding the ball to avoid defending, which is a legitimate strategy, but it is not control. Real control is measured by one question: after each touch in the opponent's half, how many seconds does that team spend in a position capable of causing damage? Croatia in 2026 had 58 percent possession but did not play control football in the traditional sense. They held the ball to control the tempo of the match, not to control the position of the ball. Those are two entirely different things merged into a single number on a stats sheet, and that merger is where data gets misread. I remember an evening in Shanghai when a young coach handed me his team's numbers after a 0-2 defeat: 63 percent possession, 612 passes, 89 percent pass accuracy. He said his team deserved to win. I asked to see the pass map. Of those 612 passes, 411 were in his own half or along the lateral corridor at halfway. Completed passes into the opponent's box: nine in ninety minutes. His team had a midfield that passed beautifully to each other and nobody forced to make a hard decision. I remember that evening because he was angry with me for about twenty minutes, then silent, then grateful. Defending has its own data gaps, and this is the kind I have cared about most throughout my career. Most defensive metrics are built around what happened: tackles, blocks, clearances, goals conceded. But elite defending is largely about what did not happen. A centre-back who reads the pass and steps up half a metre before the ball is played generates no metric at all. No tackle, no block, nothing to count. On the stats sheet he appears as a man who did nothing. On the footage he is the man who deleted an attacking intention before it took shape. Defensive data does not lie, it only falls silent when you need an answer. And it falls silent exactly at the decisive moment, because proactive defending is choosing where to fall, not where to stand still. A well-organised back line is not a wall; it is a sequence of decisions about where to concede space, where to foul, and where to collapse if things go wrong. No statistical system counts decisions that never produced an event. That is why I spend most of my viewing time on passages where nothing happens. The ball travels into a corridor, the full-back drops two metres, the holding midfielder rotates, the ball comes back out. Nobody scores, nobody tackles, nothing appears in the summary. But if I skip that passage, I will not understand why in the 70th minute that corridor suddenly widens and a goal is conceded. In 2026, when stadiums around the world closed, I got to observe a natural experiment no league would have volunteered for. The empty stadiums of 2026 did not kill football, they stripped old tactics bare. With no crowd noise, the pressure from the stands vanished, and teams that had lived off pushing referees into favourable home decisions suddenly lost a portion of an advantage they had never admitted possessing. At the same time, teams dependent on a crowd generating emotional tempo for them lost that too. Across roughly the first six to eight rounds of the empty-stadium period I logged a pattern I found notable: teams whose attacking structure rested on rehearsed, automated patterns maintained almost identical chance production; teams living on collective inspiration and crowd stimulation dropped noticeably. But I must be explicit: I logged a pattern, I did not prove a cause. The empty stadiums overlapped with fixture congestion, with injuries, with more substitutions. Anyone claiming the empty stadiums were the sole cause is filling a data gap with a story. I say that not to dismiss observation. I say it because I have seen too many confident analyses assert causation from a single variable. In this trade, confidence is usually inversely proportional to the number of variables controlled. And here is where I am heading. The biggest trap in analysis is not missing data. The trap is the reward the media system hands to whoever dares to fill a gap with a confident-sounding claim. Look at how a data report travels through the system. The analyst receives a dataset with many empty cells. He writes an internal report stating plainly that twelve metrics have no data, that conclusions rest on three matches, that reliability is low. The editor receives it and needs a headline. The headline cannot be twelve metrics without data. The headline must be something. And so the most honest report in the building is compressed into an assertion it never deserved. I have lived inside that system for forty-five years. I have covered eight Olympic Games, eight World Cups, numerous editions of the Giro d'Italia and the Tour de France. Across all those sports I have seen the same pattern: the person who makes the boldest prediction is remembered longest, the person who says I do not know yet is forgotten fastest. But across those same forty-five years I have learned something else: the people who survive longest in this trade are not the ones who predicted correctly most often, but the ones who made the fewest predictions. This is where I must talk about the transfer window, because that is the environment where data silence is exploited most thoroughly. A transfer window is really a market for buying safety for a hot seat. A deal is not merely the purchase of talent. It is the purchase of a hedge for a coach under pressure, for a club's budget, for a sporting director's reputation. When you read a transfer rumour you must read it across two balance sheets: the tactical value of the player, and the survival value of the people pushing the deal forward. That is why in a transfer window I never start from the question of how good the player is. I start elsewhere: who wins if this deal succeeds, and who loses their job if it fails? The answer to the second question usually explains the deal's momentum better than any statistic about the player. But the more important point for readers is this: in a transfer window, a club's silence does not mean the club is doing nothing. It usually means the opposite. When a club stays entirely silent while the media reports three different targets, the likely reality is that it is negotiating with a fourth nobody knows about. Noise is a tool. Silence is a tool. Readers need a filter built on structure rather than volume: how is the release clause written, how much contract remains, how much wage headroom exists, and is this player a direct replacement for one specific position. A transfer with no specific position to replace is usually a transfer not driven by tactical need. And a transfer not driven by tactical need is usually driven by somebody's survival need in the boardroom. That is not always bad. But the reader should know what they are reading. Now referees, because that is the domain where data silence has the most direct impact on results. There is a view many call a conspiracy theory, and I consider that label intellectual laziness: referees treat big clubs and small clubs differently. This requires no conspiracy. It requires only pressure. A referee working in front of seventy thousand people at a big club's home ground, in a match broadcast live to more than a hundred countries, is under a completely different cognitive load from a referee working in a lower division in front of three thousand. Higher cognitive load does not mean that referee is biased. It means his decision threshold shifts toward caution in high-controversy situations, and that shift is unevenly distributed between the two teams, because the pressure is unevenly distributed. I spent years measuring what I call decision gaps. That is the interval between two similar incidents in the same match, where one side got the decision and the other did not. If you log twelve similar contact incidents in a big match and count the whistles, you find an uneven pattern, and that pattern usually tilts toward the side with more spectators, more media pressure, or simply the side leading. I stress: this is not evidence of cheating. It is evidence that humans decide differently when cognitive load changes. That is a far less exciting conclusion than a conspiracy, and therefore it is discussed far less. VAR complicates matters in a way few anticipated. The technology removes some clear errors, but it does not remove the human decision threshold sitting in the control room. The VAR official still decides which incidents are clear enough to intervene on, and that standard is not a fixed number. It is a judgement. In big matches, the intervention bar tends to be higher, meaning minor incidents tend not to be reviewed. In smaller matches, the bar tends to be lower, meaning similar incidents tend to be reviewed. This is a paradox: more technology can produce more inconsistency across competition tiers. And once again, silence is the biggest problem. Refereeing errors are recorded, analysed, argued about. But the whistle that never comes is analysed by nobody. An incident the referee should have penalised but did not appears in no statistical table. It does not exist in the data. And because it does not exist in the data, it does not exist in the debate. Teams routinely passed over on important decisions accumulate damage across seasons, and that damage is invisible. This is the most perfect injustice in football: it leaves no trace. Let me tell one personal story to illustrate. In my first year working in China I followed a lower-division team across a full season. That team kept losing in the second half. The popular explanation was poor fitness. I cut all fourteen defeats and measured ball-in-play time. The finding had nothing to do with fitness: the team tended to be penalised roughly 40 percent more in the second half, and the rate spiked in the final twenty minutes. The foul count alone says nothing about fairness, because the team could genuinely have fouled more as it tired. But when I cross-referenced footage, most of those fouls were contests similar to incidents not penalised in the first half. I could not conclude bias. I could only conclude one certain thing: as that team tired, the referee's decision threshold toward them changed. Thirty-eight years later I still hold that conclusion and still cannot prove more. That is the nature of this trade. You produce a correct observation and it is never strong enough to become proof. You live with it. Back to the central question: when data goes silent, what must an analyst do? My answer, after forty-five years, is a four-step procedure I apply to every piece. The first step is to count the empty cells before reading the numbers. Before I look at any figure in a dataset, I check what is missing. If a sheet has twenty metrics and only twelve cells filled, I know my conclusions have a hard ceiling, and I state that ceiling in the piece. This is the most skipped step in the industry. The second step is to classify the gap into the four kinds I described: missing collection, hidden, late, misread. Each needs a different response. Missing collection needs admission. Hidden needs self-measurement. Late needs checking whether the system changed. Misread needs returning to the original footage. The third step is to find one specific passage for every conclusion. This is the discipline I impose on myself: no conclusion is allowed to survive in my writing if I cannot point to a moment on tape illustrating it. If I say a team's defence is well organised, I must point to the minute, who moved, who conceded space. Without a passage, the conclusion is deleted. The fourth step is to state plainly the part I do not know. Not as a disclaimer at the end, but as a structural part of the piece. I believe the most interesting part of an analysis is where the analyst states his own limits, because that is where readers learn to assess matches themselves instead of depending on the conclusion. This brings me to the most counter-intuitive point in the whole piece. People assume a clean report — no red flags, no issues detected — is a good report. In reality a clean report can signal two opposite things: a genuinely healthy system, or a data pipeline that has died. In football analysis these two situations look identical on paper. A team with no alarm metrics in defence might be a good defensive side, or might be a side whose data staff stopped tracking last season. A player with no injury flags in the report might be a fit player, or a player whose medical file stopped being updated. Confusing these two states produces expensive wrong decisions. I have watched a club sign a player on the strength of a spotless fitness report, only to discover the player had a documented knee issue that was never carried across in the file. The report was not wrong. It was empty. So my first rule when reading any report is: an empty state is not a safe state. An empty state is an unassessed state. The two are entirely different in consequence. For readers following Vietnamese and regional football, this deserves emphasis: leagues in Southeast Asia generally operate at far lower levels of data transparency than Europe's top leagues. This is not a moral failing of the competition. It is a feature of the terrain. Camera numbers, operating budgets, pitch quality, fixture density, travel distance between away matches — all of it creates a different data environment. An analyst applying a Bundesliga model to a league with different fixture density, different pitch quality and different travel distances is not analysing. He is copying. I learned this late. I was born in Germany, raised on the belief that tactical discipline is universal. My first six years working in Asia taught me that every league has its own gravity, and that gravity shapes everything from how a full-back stands to how a coach picks his 75th-minute substitute. There is nothing wrong with learning from European football. It is wrong to impose it without checking local conditions. What does this mean for the reader? I think it means readers need a different standard for judging analysis. A good analysis is not the one with the most certain conclusion. A good analysis is the one that shows how much data it stands on, where that data came from, and how the conclusion would change if the data were wrong. A good analysis does not cite numbers to impress. It cites numbers to narrow possibilities. A number has value only when it eliminates a hypothesis. If a number eliminates nothing, it is decoration. And a good analysis is not afraid to say there is not enough data to conclude. In this transfer window, fans will read hundreds of rumours a week. I propose a simple four-question filter. First: who is the source, and what do they gain if this rumour spreads? Second: which player does this signing replace in the XI, and if nobody, why? Third: how much contract remains, and how is the release clause structured? Fourth: if the deal collapses, who is accountable, and is that person in a position where they must succeed? Those four questions will filter most of the noise. What remains is worth tracking. For matches, I propose a different way of watching. Try once watching a game looking only at what does not happen. Watch ten minutes observing only the full-back when his team has the ball. Where does he stand, when does he move, what space does he leave behind. Then rewatch those ten minutes and count how many times the ball travelled into the space you just observed. You will learn more from those ten minutes than from a whole match watched conventionally. And you will understand why professional analysts spend most of their time on passages nobody remembers. That is the final paradox of this trade. Its entire value lies in moments that, if you do not record them, nobody will remember. People call me a tactical wizard; I just read the match one beat earlier. That beat does not come from a special gift. It comes from accepting an uncomfortable truth: most of the time, I do not know. And my job is to show others the boundary between what I know and what I do not, not to erase that boundary. So before every match I still check the empty cells in the spreadsheet before reading a single number. I still write down what I do not know at the beginning, not the end. And I still keep one rule I never break: if I have to invent a data point to finish the piece, the piece does not deserve to exist. Defensive data does not lie, it only falls silent when you need an answer. And the real test of an analyst is to stand still inside that silence rather than fill it with his own voice. Your test for the next match is this. Pick a team you follow. List five things you genuinely know about them from data you have watched yourself, not data you have read. Then list five things you assume to be true but have never verified. If you can do that, you are a long step ahead of most people who write about football. And if you discover that four of your five assumptions trace back to a single match you watched three months ago, you will understand why I say the biggest trap in this trade is not missing data, but confidence built on a gap nobody bothered to count.

The Empty Spreadsheet and the Trap of a Clean Report

Cầu thủ liên quan