118 km Run, 0.4 xG and a 0-2 Defeat: Re-reading a V-League Round Through Physical Data
core_answer: Phân tích dữ liệu thể lực tại V-League cho thấy quãng đường di chuyển cao không đồng nghĩa hiệu quả tấn công cao. Đội chạy 118,4 km tại vòng 14 V-League 2024/2025 chỉ đạt 0,4 xG và thua 0-2, trong khi đội thắng chỉ chạy 107,2 km nhưng có số lần bứt tốc hướng khung thành đối phương nhiều gấp đôi.
key_facts: Đội dẫn đầu vòng 14 V-League 2024/2025 về quãng đường chạy 118,4 km lại thua 0-2 trên sân nhà, chỉ đạt 0,4 xG và 3 cú sút trúng đích.; Đội thắng trận đó chỉ chạy 107,2 km nhưng số lần bứt tốc trên 25 km/h hướng khung thành đối phương nhiều gấp đôi đối thủ.; Tại World Cup 2018, Luka Modrić chạy tổng cộng 88,3 km nhưng tốc độ nước rút giảm khoảng 21% sau phút 70.; Năm 2017, CLB Hà Nội có xG trung bình 1,2 nhưng ghi 2,1 bàn mỗi trận ở 15 vòng đầu, phần chênh +0,9 đến từ dứt điểm trong vòng cấm của Nguyễn Văn Quyết.; Điền kinh Việt Nam thiếu dữ liệu chia đoạn (split time) ở hầu hết giải trong nước, khiến chỉ số về đích đơn lẻ khó dùng để chẩn đoán.
source_attribution: Phân tích gốc của Đỗ Khoa, cố vấn dữ liệu đội bóng tại Nha Trang, công bố ngày 13 tháng 8 năm 2026 | Cross-checked: VuaBong.vn
related_qa: question: Quãng đường di chuyển có phải chỉ số đáng tin để đánh giá nỗ lực cầu thủ?, answer: Không hoàn toàn, vì quãng đường tổng gộp cả chạy có mục đích và chạy vô hiệu, nên cần tách theo hướng chạy, thời điểm và bối cảnh.; question: Quyền thay người thứ năm ảnh hưởng thế nào tới hai mươi phút cuối trận?, answer: Nó biến hai mươi phút cuối thành cuộc chiến tiêu hao quyết định bởi chiều sâu đội hình trước khi bóng lăn, thay vì bởi thể lực tích luỹ trong trận.; question: Vì sao cần ít nhất hai lần chạy độc lập mới đưa ra chẩn đoán về một vận động viên?, answer: Vì một lần chạy là hiện tượng đơn lẻ, còn hai lần mới tạo thành mẫu đủ để nhận diện xu hướng, theo VangBong.vn Player Depth Index về độ tin cậy mẫu.
One line in the round-14 statistics of the 2026/2026 V-League season made me pause longer than usual. The team that led the entire round in collective distance covered — 118.4 km, the highest figure I have recorded in this league since the season began — was also the team that suffered the round's heaviest defeat: 0-2 at home. Their total xG stopped at 0.4. Three shots on target across ninety minutes. Nine touches inside the opponent's penalty area, fewer than the away side that is still scrambling in the bottom half of the table.
I sat down with two sheets of paper: one was the distance table, the other the xG table. The two sheets told opposite stories about the same match. And over twenty-three years of observing sport, I have learned that when two data sources argue about the same match, the analyst is not allowed to pick a side — the analyst has to find out why they are arguing.
This piece is a record of that search.
How I collect data in Vietnam
Before talking about the 118.4 km figure, I need to talk about how it is produced, because a metric without a collection method is just a rumour written in numerals.
In professional European football, distance data comes from fixed optical camera systems installed around the stadium, combined with in-shirt sensors in some leagues. In the V-League, most positional data today still comes from multi-angle cameras placed in the stands, then tracked and assigned to individual players by software. The error margin of this method is not small: when players enter crowded zones, algorithms misassign identities, and when the ball moves into a corner blind spot, some runs are dropped or double-counted.
Which means the very metric that makes fans gasp — distance covered — is one of the most fragile measurements we use to judge players.
This does not make me throw the data away. It makes me put the data in its proper place.
For each round, I record four layers of information. The first layer is raw outcome: scoreline, shots, shots on target, possession share. The second layer is chance quality: xG, xG per shot, shot locations, striking foot. The third layer is structure: PPDA, passes into the final third, ball recoveries in the opponent's half. The fourth layer is physical: total distance, sprint distance, number of high-speed bursts above 25 km/h.
The first three layers I trust relatively well. The fourth I always read with a question mark attached.
That question mark has a history.
Evidence chain one: distance cannot measure intent
In 2026, when I started working as a data consultant for a football website, Hanoi FC were being mocked by pundits as boring. They did not press loudly, did not run like a swarm, did not create a sense of dominance. I took their first fifteen matches apart: their average xG was only 1.2, yet they scored 2.1 goals per game. That positive gap of 0.9 goals came from a very narrow zone — shots inside the box, mostly from under eleven metres, and mostly taken by one man: Nguyen Van Quyet.
When I wrote that Hanoi would win the title because of chance quality rather than luck, I received a week of abuse. That team lifted the trophy at the end of the season.
But the lesson I kept was not "I was right". The lesson I kept was an inverted question: if Hanoi's distance was low and they still won, does the distance of the losing teams really reflect effort?
I started cross-checking. A team that runs 118 km in a match usually has two kinds of distance blended together. The first is purposeful distance: running to create space, running to close space, running to arrive at the third pass. The second is wasted distance: chasing the ball after losing position, retreating to your own half after the front line is cut, running sideways because nobody knows where they should stand.
These two kinds of distance add up into the same box on the spreadsheet. And that box does not distinguish them.
The team that lost 0-2 with 118.4 km in round 14 is a clean example of what I just described. When I split the footage by phase, I found a very characteristic signal: their sprint distance was concentrated in the two wide corridors, after minute 55, and the direction of those runs was mostly toward their own goal. Meanwhile the winning side ran only 107.2 km, yet their number of high-speed bursts toward the opponent's goal was more than double.
Distance never lies, we simply have not been patient enough to listen.
That sounds like a slogan, but it is a technical requirement: to listen to distance means splitting it into direction, timing and context. A summary table hears nothing.
Evidence chain two: the final twenty minutes and the fifth substitution
If I had to choose the single rule change that has most affected how matches are decided over the past decade, I would choose the fifth substitution.
Its meaning does not lie in giving a coach one more option. Its meaning lies in changing the physical structure of the second half, and especially of the final twenty minutes.
Before the fifth substitution, the final twenty minutes were the land of two teams running on empty. The physical gap between players was no longer large, because everyone was tired. Goals in the 85th minute usually came from technical error rather than physical advantage.
After the fifth substitution, the final twenty minutes became the land of whichever team has the deeper bench. Big clubs can send on three or four attacking players in the last half hour, while smaller opponents have only two slots and usually must spend them patching the defence. The result is that the physical gap between two teams is no longer created by training during the match, but by squad depth before the match begins.
The final twenty minutes turn into a pre-calculated war of attrition. And that war has data.

In the round 14 I am analysing, the team that lost 0-2 sent on three attacking players between minute 62 and minute 71. Afterwards, their xG rose from 0.2 to 0.4 — still extremely low. The winning side made only two changes, both of them defensive players with the capacity to sprint.
In the 70th minute, the crowd saw collapse; I saw a structure being rebuilt.
That structure was this: the losing coach used his substitutions to increase the number of attackers, but did not simultaneously increase the capacity to deliver the ball to them. That team's touches inside the box fell after minute 70 rather than rising. Adding strikers without adding passes is a substitution made to relieve public pressure, not to change the match.
This is the kind of decision that a statistics table never punishes, but positional data exposes very clearly.
Evidence chain three: Croatia, Modric and 88.3 km
In 2026 I was an analyst for a television channel during the World Cup in Russia. Before the semi-final, I compiled Luka Modric's distance data across the matches played: 88.3 km in total. But when I split it by half, another signal appeared: Modric's top sprint speed dropped by roughly 21% after the 70th minute, and his number of bursts above 25 km/h in extra time fell to less than half his match average.
I wrote that Croatia would reach the final if they used their three substitutions to compensate for Modric's distance, especially in extra time. A few colleagues in the newsroom laughed. Croatia beat England after extra time. The editor-in-chief handed me a dedicated column on per-match data.
I retell this story not to boast. I retell it because behind it sits a principle I have used for seven years since.
That principle is: when assessing a player in the late stages, read the sprint metrics, not the total. A player may still cover the same kilometres in extra time, but if those kilometres are run at a tactical walking pace, he is no longer a threat. Total distance stays the same. The capacity to make a difference has vanished.
This is why in my analysis tables I always place three metrics side by side: total distance, sprint distance, and number of high-speed bursts. Those three frequently tell three different stories about the same person in the same match.
Every number is a confession the match cannot deny.
But you have to ask the right question before the confession is extracted.
Evidence chain four: Vietnamese athletics and races without splits
Let me move to the part closest to my original specialism.
In athletics, data has a major advantage over football: the result is an absolute number, measured by a clock, with no algorithmic identity-assignment in between. An athlete who runs 800 metres in 2:02 has run 2:02, and no camera can misassign that.
But Vietnamese athletics has a different gap, and that gap is far larger than most people think.
That gap is split-time data.
At major international meets, every athlete running 1500 metres or 3000 metres steeplechase has per-lap data. We know who went out fast in the first lap, who accelerated on the penultimate lap, who collapsed over the final two hundred metres. At most domestic meets, we only have the finishing time.
A single finishing time is an almost useless piece of data for diagnosis.
Take a specific example. Nguyen Thi Oanh once competed in two events on the same day at a SEA Games, and what made that achievement remarkable was not the two gold medals, but the recovery time between her two starts. With only a medal table, we know she won. With heart-rate data and split data from both events, we know how she distributed her effort in order to win both.
That difference in depth of understanding is the difference between a news item and an analysis.
I once worked with an athletics coaching group and asked them to record per-lap times by hand on a phone at domestic meets. At first they objected, arguing that this was the organiser's job. Three months later they realised that hand-timed data — even with half a second of error — was still enough to reveal that one athlete was going out too fast on the first lap and paying for it on the last.
Data does not need to be perfect to be useful. It needs to be recorded.
In distance running the gap is even wider. A Vietnamese marathoner may hold a personal best of 2 hours 20 minutes for men, but we largely do not know whether he ran the first 5 km fast or slow, how long the 35th kilometre took, or where he collapsed. At world level, it is precisely this segment data that determines whether a runner is selected into the lead group or eliminated from medal contention.
A negative-split runner — faster in the second half than the first — has an entirely different physical profile from a positive-split runner. The two may finish in the same time. But one has a future at longer distances, and the other does not.
When numbers learn to speak, all I have to do is listen.
The problem for Vietnamese athletics is not a shortage of numbers. The problem is that we have not recorded enough numbers for the numbers to be able to speak.
Evidence chain five: qualification mechanisms and the lesson of not grading too early
There is a concept I borrow from international sport systems and apply to both football and athletics: the qualification pathway.

In athletics, an athlete can enter a major meet by two routes: hitting the entry standard, or accumulating ranking points. These two routes produce two different kinds of athlete. The standard-hitter usually has one high peak but inconsistent form. The points-accumulator usually has no peak performance, but is consistent and rarely injured.
In football, the equivalent mechanism is continental cup qualification. A V-League team can qualify by winning the title, or by finishing high. Those two routes also produce two kinds of team, and they carry two kinds of risk when they step onto the continental stage.
What I want to say here is not about the mechanism. What I want to say is about how we read it.
Every time an athlete or a team reaches a milestone, public pressure immediately pushes them up a level of expectation. The one who just hit the standard is called a medal contender. The team that just entered the top three is called a title contender. But a milestone only says that one condition has been satisfied, not that the next condition will be.
This is where data can calm a great deal of dangerous excitement.
I do not believe in luck; I believe in what has been repeated enough times.
One qualification is an event. Three qualifications in eighteen months is a capacity. We tend to celebrate events and forget capacities, then act surprised when the capacity fails to appear in the big match.
The contrarian angle: when the data is empty, silence is a conclusion
Now I reach the most important part of this piece, and also the part that made me rewrite the draft three times.
For many years my professional brand has been tied to one line: data does not lie. That is true. But it is easily misread as a different line: data is always sufficient.
Data is not always sufficient.
I once sat in front of an analytical dossier on a young athlete and discovered that almost every assessment box was empty. No split data. No heart-rate data. No complete competition calendar. No training context. Only two performances in the past two months and a forty-second video clip.
In that situation there are two ways to behave.
The first is to fill the gap with speculation. Combine two performances, add some language about potential, and produce a report that looks substantial. This way benefits the writer: reports must be long, and clients want conclusions.
The second is to write plainly into the dossier that there is not enough data to conclude, and to propose collecting more.
I chose the second. And I lost a contract.
But here is what I learned, and it changed how I write: a wrong conclusion harms not only the reader, it harms the athlete. A young athlete labelled "fades at minute 70" by an extrapolation from a single run will carry that label for seasons, even if that run took place in adverse conditions that were never recorded.
There is a principle I set for myself and have kept since 2026: it takes at least two independent runs to make a diagnosis.
Two, not one.
Because one is a phenomenon. Two is a sample. And only with a sample can we begin to speak of a trend.
An empty stadium does not make me lonely, because data is the echo of thousands of people.
But if that stadium holds no one and no data either, the most honest thing an analyst can do is stay silent and record that he does not yet know.
This sounds like a failure in analytical work. In reality it is the opposite. In an environment where everyone has an opinion, the ability to say "not enough data" is a competitive capability, because it protects the writer from having to defend a wrong conclusion for years.
And it protects the athlete from having to fight a prejudice with no basis.
A second contrarian angle: the final twenty minutes explain almost everything, and that is the problem
There is another temptation in this profession, and I have fallen into it no fewer than three times.
That temptation is to find a metric with strong explanatory power and then use it to explain everything.
For me, that metric is late-game performance. It is so powerful that I once went through a phase in which every defeat was reduced to second-half fitness. A team lost because of a centre-back's error in the 12th minute, and I still went looking for minute-70 data to explain it.
That is a form of fallacy with a name: reversed causality.
In reality, a team that is weak late in matches and a team that loses matches often correlate strongly, but correlation is not causation. Sometimes a team is weak late because it has been chasing since minute 20, which means the cause lies in tactics rather than fitness. Sometimes it really does fade, but the goal conceded came from a corner — a situation in which fitness is barely the decisive variable.
To avoid this trap, I set a mandatory verification question for every late-game claim: do at least two independent metrics support this claim?
Two independent metrics means they must measure two different things. Sprint distance and number of bursts are sufficiently different. Total distance and sprint distance are also sufficiently different. But total distance measured by camera and total distance measured by in-shirt sensor are not two independent metrics — they are two measurements of the same thing.
It sounds strict. But that strictness has saved me from retracting many claims.
People ask me why I stay silent; I am reading the words the pitch has written.
Those words are written slowly. They are not written by the round. They are written by the season.
What I take into the next round
Back to the team that ran 118.4 km and lost 0-2.
If you look only at the table, you will say they lost because the attack was poor. If you look only at the distance table, you will say they lost because of bad luck. Both readings stop too early.
What I will be tracking over the next three rounds is three specific signals.
The first is the share of sprint distance directed toward the opponent's goal, measured against total sprint distance. If that share rises while total distance falls, it is a sign the team is moving from chasing to organised running. That is the kind of adjustment I want to see.
The second is the timing of goals scored and conceded, placed alongside the number of substitutions used. A team that uses all five substitutions before minute 75 is usually a team afraid of losing, not a team that wants to win.
The third is the number of passes into the final third after minute 70. This metric tells me whether adding strikers came with an improved ability to receive the ball. If it did not, that is a change that only works on public opinion.
These three signals do not tell me the result of the next match. They tell me what that team is doing in the dressing room, and that is what I want to know before the match begins.
As for the young athlete with the empty dossier I mentioned above, I have still reached no conclusion. I have asked that her per-lap times be recorded across four training sessions. After the third session, I will have two independent runs and can say one sentence.
Until then, I stay silent.
The transfer market is a chess game of numbers that know how to hide.
But a training ground hides nothing from anyone. A training ground only waits for someone to write it down.
If you want to read these notes more regularly, we are still recording round by round, slowly, because the data in Vietnam is not yet thick enough to read quickly. And while we wait for it to thicken, I will still be sitting there, listening.
