12 models evaluated on Vietnamese real estate price prediction using repeated stratified cross-validation on 81,790+ listings.
CatBoost dẫn đầu ở RMSE 0.336
Tuyển chọn mô hình 04/2026 (5-fold CV, mô hình một đầu ra) — độ chính xác hiện tại nằm ở thẻ Production drift.
default
default
Pre-trained on 130M synthetic datasets, no fine-tuning
catboost-3quantile-v1-20260831 · out-of-sample: only listings first seen after this model's training cutoff, so the error is un-memorisedCoverage materially below 80% means the band is — it is narrower than the model's actual error. Unequal tails mean it is also : a high “below p10” share means the model prices listings above what sellers actually ask. Measured on 31 interval-bearing predictions. This is a known open defect, not a validated interval.
overall_oos), cửa sổ 7 ngày không chồng lấn, gần nhất kết thúc 2026-09-04.Hôm nay mô hình chưa đạt 4/4 điều kiện. Đây là kết quả dự kiến: sai số ngoài mẫu hiện 37.5% so với ngưỡng 15%, và khoảng p10–p90 phủ 64.5% so với mức danh nghĩa 80%. Ngưỡng được viết ra trước chính là để khoảng cách hiện rõ thay vì bị định nghĩa lại cho vừa. Thẻ này là thước đo, không phải bảng điểm.
Độ phủ p10–p90 ngoài mẫu nằm trong [78, 82]% và mỗi đuôi trong [8, 12]%, giữ được 4 cửa sổ tuần liên tiếp không chồng lấn.
Hiện tại: 0/4 tuần nằm trong dải · 1/4 tuần có đo hiệu chuẩn · gần nhất phủ 64.5% (dưới p10 12.9%, trên p90 22.6%)
| Endpoint | State | Requests | p50 | p95 | p99 | Error rate |
|---|---|---|---|---|---|---|
| fast | warm | 55 | 1307 ms | 1757 ms | 1927 ms | 29.09% |
Mô hình dự đoán log của đơn giá (giá mỗi m²), nên RMSE và MAE nằm ở thang log — không phải phần trăm. Muốn biết độ chính xác theo phần trăm, hãy nhìn vào MdAPE.
Đây không phải độ chính xác người dùng nhận được hôm nay (một lượt so sánh họ mô hình chạy một lần vào 04/2026 trên 81,790 dòng đã lọc, bằng cross-validation, với bộ dự đoán một đầu ra đã được thay bằng ensemble 3-quantile đang chạy): con số nên trích dẫn là MdAPE ngoài mẫu ở thẻ Production drift phía trên (~37%).
| # | Model | Config | RMSE | MAE | MAPE | Time/fold | Rows |
|---|---|---|---|---|---|---|---|
| 1 | CatBoostdeployed | default | 0.3360 | 0.2246 | 25.2% | 24s | 81,790 |
| 2 |
A neural network pre-trained on massive data to serve as a reusable base for many downstream tasks. Unlike traditional models that train from scratch on your specific dataset, a foundation model arrives already knowing general patterns — you just show it a handful of examples and it generalizes. This is the paradigm behind GPT (text), CLIP (images), and now TabICL (tables). Pre-training on 130M synthetic tabular datasets lets TabICL do in-context learning: predicting on new data without any fine-tuning.
Foundation model · 130M synthetic datasets
A transformer pre-trained on 130 million synthetic tabular datasets, learning the statistical patterns common across all tabular data. At inference it performs in-context learning: you show it a handful of labeled examples plus a test row, and it predicts — without any training on your specific dataset. The tabular equivalent of GPT's few-shot prompting. Wins statistically significantly at small scale (<1K rows); gradient boosters catch up as data grows. Runs on an A10G GPU, ~50s latency.
The leaderboard is stale and describes a different model. Its cross-validation figures (best CatBoost MdAPE 14.8%, 81,790 rows, 2026-04-15) come from a single-output point predictor that predates the 3-quantile ensemble now deployed, and regenerating it is a manual step that has not run since. Treat it as a model-family comparison, not as an accuracy claim.
The figure to quote is the Production drift card above: median absolute percentage error on live, out-of-sample scoring — listings first seen after the model's training cutoff. That currently runs around 30%, roughly twice the cross-validation number, because CV measures the filtered training distribution while production scores every active listing.
Currently deployed: CatBoost (depth=10, autoresearch-tuned) for fast inference. TabICL (foundation model) available as an alternative.
Last updated: April 16, 2026
MdAPE ngoài mẫu ≤ 15% toàn cục và ≤ 18% cho từng loại hình, với n ≥ 1.000 dự đoán đứng sau mỗi con số.
Hiện tại: toàn cục 37.5% trên n=31 · 0/4 loại hình chưa đạt
Không cửa sổ nào trong 8 tuần gần nhất vượt 20% MdAPE ngoài mẫu.
Hiện tại: 8/8 cửa sổ vượt ngưỡng · cao nhất 37.5%
Công bố MdAPE ngoài mẫu theo từng tỉnh/thành; thí điểm chỉ mở ở nơi đạt ngưỡng, không mở theo trung bình toàn quốc.
Hiện tại: chưa có dòng province_oos nào trong prediction_drift
Thí điểm chỉ mở ở từng tỉnh/thành đạt ngưỡng, không mở theo trung bình toàn quốc. Một tỉnh đủ điều kiện khi MdAPE ngoài mẫu ≤ 15% trên ít nhất 1.000 dự đoán trong cửa sổ — cùng một tiêu chuẩn với con số toàn cục, không hạ thấp.
Chưa có dữ liệu tỉnh — sẽ xuất hiện sau kỳ drift tuần tới. Chiều province_oos vừa được thêm vào compute_drift.py; bảng này tự điền khi tác vụ drift chạy lần kế tiếp trên môi trường sản xuất.
Nhãn là giá rao, không phải giá giao dịch. Đạt ngưỡng này chứng minh mô hình khớp với mức giá người bán rao, không phải mức giá tài sản thực sự được sang tay. Cầu nối giá rao → giá giao dịch (proxy từ tin bị gỡ) sẽ được đo, chưa đưa vào sản phẩm, trong năm đầu. Định nghĩa ngưỡng dưới dạng mã nằm ở src/ml/pilot_bar.py.
Bảng xếp hạng bên dưới dùng cross-validation; thẻ Production drift đo trực tiếp trên dự đoán sản xuất — đó mới là con số MdAPE nên trích dẫn (~37%).
| lr=0.030 |
0.3816 |
| 0.2574 |
| 29.4% |
| 35s |
| 81,790 |
| 3 | XGBoost | default | 0.3836 | 0.2630 | 29.9% | 20s | 81,790 |
| 4 | AutoGluon | best quality, 600s limit | 0.3993 | 0.2717 | 30.3% | 12m 20s | 90,566 |
| 5 | TabICLfoundation | n=32 | 0.4995 | 0.3256 | 38.9% | 1m 59s | 90,566 |
| 6 | KNN | default | 0.5180 | 0.3450 | 42.1% | 29s | 90,566 |
| 7 | Ridge | default | 0.5516 | 0.3644 | 43.5% | <1s | 12,908 |
| 8 | RandomForest | default | 0.5640 | 0.3920 | 46.9% | 18m 4s | 90,566 |
| 9 | LinearRegression | default | 0.5779 | 0.3781 | 46.8% | 37s | 90,566 |
| 10 | LinearSVR | default | 0.5922 | 0.3638 | 43.9% | 23s | 90,566 |
| 11 | ElasticNet | default | 0.6102 | 0.4144 | 49.4% | 15s | 90,566 |
| 12 | Lasso | default | 0.6196 | 0.4224 | 51.0% | 8s | 90,566 |
Amazon · stacked ensemble AutoML
Not a single model but a system. Trains LightGBM + CatBoost + XGBoost + Random Forest + neural nets + FT-Transformer in parallel, then stacks them with a meta-learner. The tradeoff for its accuracy is 10-minute fold times and a harder-to-deploy artifact.
Yandex · native categorical handling
Gradient boosting that builds decision trees sequentially — each new tree corrects the previous tree's errors. CatBoost's edge is native support for categorical variables via ordered target statistics, which matters enormously here: we have 257 districts, 235 wards, and 974 streets that other frameworks either one-hot-encode (explosion) or ordinal-encode (lossy). Deployed in production at depth=10, autoresearch-tuned.
Microsoft · leaf-wise, histogram-based
Microsoft's speed-optimized gradient booster. Uses histogram-based feature binning and leaf-wise tree growth (vs XGBoost's level-wise), which converges faster on most tabular problems. Trained here with Huber loss to resist the long tail of outlier listings, and handles NaN values natively.
The original boosting framework
Popularized gradient boosting on tabular data and still a strong baseline. Slightly behind CatBoost and LightGBM on this dataset because it needs ordinal encoding for high-cardinality categoricals (its native categorical mode fails on district-level granularity).
Bagged decision trees
An ensemble of decision trees trained in parallel on bootstrapped samples, with predictions averaged across trees. Interpretable and robust, but lacks the sequential error-correction mechanism of boosting — which is why it plateaus around RMSE 0.56 here while boosters reach 0.34.
k-nearest neighbors · non-parametric
Predicts by averaging the 10 nearest listings in feature space, distance-weighted. No training phase; pays the cost at inference time. Struggles as categorical cardinality and feature dimensionality grow.
L2-regularized linear regression
Linear regression with L2 (squared coefficient) penalty — shrinks all coefficients toward zero. Useful as a sanity-check baseline: quantifies how much of the variance is linearly explained before we invoke non-linear models.
L1-regularized · sparse
Linear regression with L1 (absolute coefficient) penalty. Drives unimportant coefficients to exactly zero, giving implicit feature selection.
L1 + L2 blend
Blends Lasso (L1) and Ridge (L2) penalties. Handles correlated features more gracefully than pure Lasso (which arbitrarily picks one of a correlated pair).
ε-insensitive loss
Support vector regression with a linear kernel. Optimizes an ε-tube (ignores errors below ε, penalizes errors outside) rather than squared error — different loss geometry, comparable expressiveness to Ridge.
Plain OLS · no regularization
Classical ordinary least squares. The null hypothesis against which every more complex model must justify its complexity.
Accuracy is measured against asking prices, not sale prices. Every label in this system is a listed price scraped from a portal. Vietnamese resale prices are commonly negotiated down from ask, so the model predicts what a property will be listed at, not what it will sell for. No transaction-price data is used anywhere.
The p10–p90 interval is measurably overconfident. The band comes from three CatBoost models trained at α=0.1/0.5/0.9. Coverage was measured for the first time on 2026-08-30 and is reported live in the calibration panel above. On out-of-sample listings it covered ~62% of actual prices against a nominal 80%, with ~22% falling below p10 against a nominal 10% — so the band is both too narrow and skewed toward over-pricing. Treat it as indicative. Promotion still selects on median-quantile MdAPE only, so nothing in the training loop currently corrects this.