Splits and Stroke Rate: Decoding Vietnamese Swimming Speed Through Raw Data
**Câu trả lời cốt lõi:** Phân tích chia đoạn trong bơi lội dùng năm trục — độ lệch thời gian theo từng 25 mét, tần số quạt tay, quãng đường mỗi lượt quạt, hiệu suất xoay người cùng pha dưới nước, và độ ổn định nhịp cuối — để dự báo thành tích tốt hơn thời gian chung cuộc. **Sự kiện then chốt:** - Độ dài phân bố nhịp (chênh lệch đoạn nhanh nhất và chậm nhất) dự báo thành tích cuối mùa tốt hơn thời gian chung cuộc; ngưỡng tốt là 4–5 giây cho cự ly 200 mét. - Nhóm đạp cá heo từ bốn lần trở lên nhanh hơn nhóm đạp ba lần trở xuống trung bình 0,42 giây ở 15 mét đầu sau khi rời tường. - Độ ổn định nhịp cuối dưới 95% cho thấy vận động viên đang bơi bằng nợ năng lượng thay vì năng lực. - Thời gian phản xạ xuất phát tương quan mạnh ở cự ly 50 mét nhưng yếu ở cự ly 200 mét trở lên. **Nguồn:** Phân tích dữ liệu theo dõi nhiều mùa của Huang Chengyu, bảng chia đoạn dựng thủ công từ video giải bơi quốc gia, giai đoạn 2022–2024. **Hỏi đáp liên quan:** - Hỏi: Chỉ số nào dự báo thành tích bơi tốt nhất? Đáp: Độ ổn định nhịp cuối và độ dài phân bố nhịp, theo dữ liệu nhiều mùa. - Hỏi: Pha xoay người đóng góp bao nhiêu vào kết quả? Đáp: Có thể lên tới gần 4 giây trong nội dung 200 mét nếu hiệu suất xoay người kém. - Hỏi: Vì sao bơi lội Việt Nam thiếu dữ liệu? Đáp: Do thiếu thiết bị chia đoạn, thiếu nhân lực phân tích, và thiếu động lực công bố dữ liệu chi tiết.
Splits and Stroke Rate: Decoding Vietnamese Swimming Speed Through Raw Data
A 0.8-second gap that was not in the final 50 metres
On the night of the men's 200m freestyle final at a national meet, I sat in the sixth row. No split screen, no official 25-metre data. I did what I have done for more than twenty years in this job: pressed a stopwatch on every wall touch, counted strokes in each segment, and noted the moment the rhythm began to change.

The final result: two swimmers separated by 0.8 seconds. When I rebuilt the split curve in a spreadsheet, that gap did not appear in the last 50 metres. It appeared in the 150-metre segment.
The winner's stroke rate fell from 46 strokes per minute in the fifth segment to 41 in the seventh. The runner-up held a steady 44 across the final two segments. The winner still won — but he won on the foundation built over the first 100 metres, not on the finishing surge. Had the race been 25 metres longer, the result would have flipped.
I never read swimming results off a leaderboard. A leaderboard answers who touched first. It does not answer who was swimming faster at the moment that mattered most.
Context: A sport measured with too few numbers
Swimming has the greatest data advantage of any Olympic sport. The environment is controlled, lanes are fixed, times are measured in hundredths, and every athlete performs exactly the same workload. No wind, no pitch, no direct contact with opponents. In theory, this should be a paradise for analytics.
In Vietnam the reality is the opposite. Most domestic meets publish only the final time and the cumulative splits at each 50 metres. Stroke rate, distance per stroke, turn time, number of underwater dolphin kicks, and reaction time — four of the five decisive metrics — are almost never recorded. A swimming nation that can produce champions but cannot explain why they won will forever depend on inspiration.
Based on my experience tracking races and meets across many seasons, I built a small five-axis framework: (1) deviation between actual and model times per 25 metres, (2) stroke rate, (3) distance per stroke, (4) turn efficiency and underwater phase, (5) stability of the closing rhythm. With these five axes I can reconstruct almost the entire story of a race without specialised equipment.
The framework is not perfect. It depends on the eye pressing the stopwatch — mine. My error range is roughly plus or minus 0.15 seconds per wall touch. But the principle holds: when you measure the same thing the same way across seasons, error becomes a constant, and constants can be subtracted. That is why I set my significance threshold before seeing results, not after. The threshold is fixed from the prior season and is not allowed to move.
Why Vietnamese swimming lacks data
There are three structural causes. First, equipment. An automatic split system for eight lanes costs more than most domestic pools can afford, and even where it exists, per-meet operating costs sit outside an organiser's budget. Second, trained people who can read data. Anyone can press a stopwatch; rebuilding a split model and comparing it to a reference model requires method. Third, and most important, incentive. If a results board is enough to award medals and name champions, nobody has a reason to dig deeper.
These causes feed each other. No equipment leads to no data; no data leads to no expertise; no expertise leads to a generation of coaches making decisions by feel, and feel cannot be verified. When a decision cannot be verified, it cannot be wrong. And when it cannot be wrong, it is never fixed.
While waiting for infrastructure, the only way to produce data is to do it by hand. That is the path I chose, and the path I recommend to young coaches. A phone filming from a fixed angle, a spreadsheet, and three patient seasons can produce a dataset no centre in Southeast Asia owns.
Core: An evidence chain from one race
Take a concrete example from data I collected over the last two seasons in the men's 200m freestyle. I choose this event because it exposes the boundary between speed and endurance most clearly. Two hundred metres sits at the crossover of two energy systems, where every mistake is punished but the distance is not yet long enough to blame pure conditioning.
The model split table for a junior national-standard swimmer looks like this, in seconds per 50 metres:
| Segment | Model | Actual | Deviation | |---------|-------|--------|-----------| | 0–50 | 27.2 | 26.5 | −0.7 | | 50–100 | 29.4 | 29.1 | −0.3 | | 100–150 | 30.6 | 31.8 | +1.2 | | 150–200 | 31.2 | 32.9 | +1.7 |
Three things stand out. First, the negative deviation in the opening two segments is not a good sign. A swimmer 0.7 seconds faster than model in the first segment alone has spent reserve energy on the opening rhythm instead of saving it for the finish. Second, deviation rises steadily after 100 metres, showing the decline is structural, not random. Third, the gap between fastest and slowest segment is 6.4 seconds. For a swimmer with good pacing distribution, that number should sit between 4 and 5 seconds.
Pacing distribution length — the gap between fastest and slowest segment — predicts end-of-season performance better than the final time itself. A swimmer 2 seconds slower but with stable pacing distribution will typically overtake a faster swimmer with skewed distribution within three to six months. I verified this on a sample of 34 junior swimmers over three seasons, and the hit rate was 71%. Not an absolute number, but far above guessing from current results.
This leads to the second axis: stroke rate and distance per stroke. The two are systematically inverse. Raising one usually lowers the other, and their product approximates swimming speed. A swimmer can go faster by stroking more often, longer, or both, but each choice demands a different kind of conditioning and technique — and pays a price at a different point in the race.
In my tracked data, swimmer A averages 44 strokes per minute at 2.05 metres per stroke. Swimmer B strokes 48 times at 1.88 metres. Their speed products are nearly identical. Two swimmers travel at the same speed along completely different paths. But the energy cost differs, and that is where the story is decided.
At the 150-metre segment, A holds 1.98 metres per stroke even as rate drops to 42. B holds rate at 48 but distance falls to 1.72 metres. The long-stroking swimmer loses less speed when tired because he has technical headroom to compensate with rhythm. The fast-stroking swimmer has no such headroom. As muscles fatigue, strokes shorten and speed collapses faster. This is why I always tell young coaches to read distance per stroke in the final three quarters of a race, not the opening. Everyone is long at the start. The finish exposes real technique.
The third axis is turns and the underwater phase. This is the most undervalued part of Vietnamese swimming. In butterfly and freestyle, an inefficient turn can cost 0.3 to 0.6 seconds each. Multiply by seven turns in 200 metres and the accumulated loss approaches 4 seconds — larger than the entire winning margin at most domestic meets.
I measure the turn as the time from head touching the wall to head leaving the wall on the first stroke. At national level, my reference thresholds are 1.1 seconds for freestyle and butterfly, 1.3 seconds for backstroke and breaststroke. Beyond these thresholds, every extra 0.1 second is equivalent to swimming 0.1 second slower over 25 metres at peak speed — a double loss, because you lose both time and momentum.
The number of underwater dolphin kicks is the second metric on this axis. International research has long shown that dolphin kicking is faster than surface swimming in the early phase of a race, and the world's top swimmers take four to six kicks before surfacing. Domestically, most swimmers take only two or three and surface too early, giving away free speed. No rule forbids more. No coach penalises more. Yet they still surface early.
I tested this hypothesis by comparing two groups in my dataset: those taking four or more kicks and those taking three or fewer. The average time difference over the first 15 metres after leaving the wall was 0.42 seconds in favour of the higher-kick group. Multiplied across the turns in a 200m event, the cumulative gap approaches 3 seconds.
Three seconds. At a national meet where the gap between gold and bronze is often under 1.5 seconds, three seconds is another world. And those three seconds require no extra hour of training. They require only a change in how you count.
The fourth axis is closing-rhythm stability. I estimate this as the ratio between average speed in the final two segments and average speed in the middle two. My sustainability threshold is 95%. Below it, the swimmer is swimming on energy debt, not capacity. Above it, there is headroom to accelerate at the finish — and that headroom, not the final time, is what sustains a career.
Another finding from the data that I consider important: reaction time correlates very weakly with final result in events of 200 metres and above, but correlates strongly at 50 metres. This sounds obvious, but domestic training programmes often spend a great deal of time on start drills across every distance while ignoring the turn — which contributes many times more at middle and long distances. This is a misallocation of resources, and the data says so clearly.
Let me build a summary comparison table for the two distance groups. This is the table I use when advising programme allocation:
| Metric | 50m events | 200m and above | |--------|------------|----------------| | Reaction time | High weight | Low weight | | Stroke rate and distance per stroke | High weight | Medium weight | | Turns and underwater phase | Low weight | High weight | | Closing-rhythm stability | Low weight | Highest weight |
This table explains why one swimmer can win 50 metres but fail completely at 200, and vice versa. Two different physiological problems, requiring two different training structures. Training them with the same programme is a systematic waste, and it happens everywhere I have been.
400m individual medley: where data exposes most
If I had to pick one event to demonstrate the power of split analysis, I would pick the 400m individual medley. This is four strokes in succession, and each segment has its own resource-allocation model. There is nowhere to hide a mistake.
At national level, the reference model I use for a junior male swimmer is: butterfly 28%, backstroke 29%, breaststroke 22%, freestyle 21%. Notably, breaststroke carries the lowest percentage weight even though it is usually the slowest stroke. That is because strong swimmers switch to breaststroke at low rate but long distance per stroke, preserving glide. Weak swimmers switch to breaststroke at high rate, short distance, and lose the entire accumulated lead from the first two segments.
In three seasons of my data, there were eleven cases of a swimmer leading after 200 metres and then losing in the breaststroke leg. All eleven shared one data signature: their distance per stroke in the butterfly and backstroke legs was above average, but their breaststroke distance per stroke was 12% or more below average. They used powerful arm strokes to compensate for weak breaststroke kick. In the final leg, that compensation ran out.
This is the kind of finding a final results board can never provide. It only emerges when you separate the four strokes and measure each independently. And it has practical value: a coach who learns the problem is in the breaststroke kick rather than general conditioning can save an entire training cycle.
Counter-intuitive angle: peak speed is the most celebrated trap
In every swimming commentary, people love the moments of acceleration. A blazing 25-metre segment, a sprint from fourth to first, a peak speed captured on camera. Those moments are real. But they are systematically misinterpreted.
Data never lies, but it knows how to hide. Consider this example from my dataset: a swimmer recorded the fastest 25-metre segment of the entire meet at 2.18 metres per second, yet his 200m final time was 1.4 seconds behind the winner. Analysis showed that peak speed came at the 75-metre mark, not the finish. He burned his reserve at a point that did not matter.
This is what split data forces you to face: peak speed and final time are two metrics that can be decoupled, and in most cases, peak speed at any given point is actually a warning sign of poor resource allocation.
People look at the results board; I look at the curve. The curve tells me what a swimmer spent, where he spent it, and how much remains. A good curve is nearly flat over the first two thirds and rises over the final third. A bad curve has an early sharp peak and then collapses.
The problem is that in swimming, prizes and sponsorship profiles often reward that peak moment — a viral clip, a highlight reel — without checking where it sits on the curve. And so we are paying for an illusion. A swimmer learns that the reward comes from producing a beautiful moment, not from swimming faster across the whole race. That is a wrong lesson, and it has derailed many young careers.
A second example, this time on the correlation between personal bests and stability. In one season I tracked two swimmers of the same age group. Both broke personal bests twice. Swimmer C broke his in two races with clean pacing distribution and closing-rhythm stability of 96%. Swimmer D broke his in two races with closing stability of 88% and 87%. The next season, C improved by another 1.1 seconds. D plateaued and suffered a shoulder injury. Same immediate result, two opposite trajectories.
Correlation is not causation. A good time does not automatically mean a good foundation. And this is the point most domestic analysis misses when it publishes only final times.
Luck is something I do not have. I have probability and thick enough data. When someone says a swimmer is finding form, I open the spreadsheet and check whether that lies in the time, in the stability, or merely in a random peak. In three cases out of four, form is just a small data sample viewed from a flattering angle.
I also have to audit myself. I once publicly predicted a young swimmer would plateau after his closing-rhythm stability fell to 90%. He improved by another 0.9 seconds the following season. My data was right on the metric but wrong on the forecast — because he changed coaches and completely restructured his programme. I logged that miss in the same spreadsheet as my correct calls. An analyst who does not publish his misses is an analyst selling an illusion.
Next cycle: the resource problem
If split data can find three free seconds in the turn phase, the question is no longer whether to do it. The question is who will do it first.
Over the next two to three months, the development to watch in Vietnamese swimming will not be a new personal best. It will be how many coaches begin recording turn time and dolphin-kick count instead of only final times. The first signal I will look for is a split dataset published openly by a domestic training centre. If that happens, the entire analytical baseline shifts within a year.
When I see a young swimmer improve closing-rhythm stability from 92% to 96% while the final time barely moves, that is a real growth signal. The time will come later. The curve moves ahead of the results board. And when a swimmer produces a spectacular surge at any given segment, I do not ask how fast he was. I ask what he paid for that segment, and what he paid with in the segment after.
A team does not collapse in one night. It collapses when the metrics stop connecting to each other. In swimming, the inverse is also true: a champion is not made in one night. Champions are made when the metrics start aligning, and that is the only thing I trust.
