From 579ea6bf85b1c8d0179f558c5db85fa3bef993e0 Mon Sep 17 00:00:00 2001 From: Vinta Chen Date: Sun, 27 Sep 2026 03:03:12 +0800 Subject: [PATCH] docs: add machine learning category intro The page had no intro, so readers got no pick among scikit-learn, the boosting libraries, pgmpy, Feature-engine, and TimesFM. Co-Authored-By: Claude --- .../data/category_intros/machine-learning.md | 29 +++++++++++++++++++ 1 file changed, 29 insertions(+) create mode 100644 website/data/category_intros/machine-learning.md diff --git a/website/data/category_intros/machine-learning.md b/website/data/category_intros/machine-learning.md new file mode 100644 index 00000000..586bbbc5 --- /dev/null +++ b/website/data/category_intros/machine-learning.md @@ -0,0 +1,29 @@ +Tabular data goes to scikit-learn, and boosted trees to LightGBM or CatBoost. Both fit in that Python machine learning library's Pipeline. + +How to choose: + +- Classification, regression, and clustering on tabular data: scikit-learn +- Boosted trees that train fast on large datasets: LightGBM +- Boosted trees on data with many categorical columns: CatBoost +- Bayesian networks and causal models: pgmpy +- Feature engineering on pandas dataframes: Feature-engine +- Boosted trees on a Spark, Dask, or Ray cluster: XGBoost +- Forecasting time series without training a model first: TimesFM + +scikit-learn covers [supervised and unsupervised learning](https://scikit-learn.org/stable/getting_started.html), plus preprocessing, model selection, and evaluation. Deep learning is [out of its scope](https://scikit-learn.org/stable/faq.html#why-is-there-no-support-for-deep-or-reinforcement-learning-will-there-be-such-support-in-the-future). To pick a model, follow its [Choosing the right estimator](https://scikit-learn.org/stable/machine_learning_map.html) chart. Split your data into train and test sets [before any preprocessing](https://scikit-learn.org/stable/common_pitfalls.html#how-to-avoid-data-leakage). Then put the preprocessing and the model in one Pipeline, so cross-validation and search never fit on the data they score. For gradient boosting without another dependency, its HistGradientBoostingClassifier and HistGradientBoostingRegressor [handle missing values and categorical data](https://scikit-learn.org/stable/modules/ensemble.html#gradient-boosted-trees) with no preprocessing. + +LightGBM aims at [faster training and lower memory use](https://lightgbm.readthedocs.io/en/latest/). It grows trees leaf-wise, so `num_leaves` is [the main parameter to tune](https://lightgbm.readthedocs.io/en/latest/Parameters-Tuning.html#tune-parameters-for-the-leaf-wise-best-first-tree): keep it below 2^(max_depth). To prevent overfitting, raise `min_data_in_leaf`. Instead of one-hot encoding, mark categorical columns with `categorical_feature`: it [often performs better](https://lightgbm.readthedocs.io/en/latest/Advanced-Topics.html#categorical-feature-support). With a validation set, use [early stopping](https://lightgbm.readthedocs.io/en/latest/Python-Intro.html#early-stopping) to find the number of boosting rounds. + +CatBoost takes [non-numeric features without preprocessing](https://catboost.ai/) and gives good results with its default parameters. Its docs say [not to one-hot encode](https://catboost.ai/docs/en/features/categorical-features) during preprocessing: list the categorical columns in `cat_features` instead. Before tuning anything else, [rule out underfitting and overfitting](https://catboost.ai/docs/en/concepts/parameter-tuning): set a large number of iterations, and turn on the overfitting detector and the use-best-model option. The learning rate is set from your data by default. + +pgmpy does [causal and probabilistic reasoning with graphical models](https://pgmpy.org/), from learning a graph from data to running inference on the fitted model. scikit-learn [leaves graphical models out](https://scikit-learn.org/stable/faq.html#will-you-add-graphical-models-or-sequence-prediction-to-scikit-learn), so use pgmpy for Bayesian networks. Every discovery algorithm [follows one pattern](https://pgmpy.org/guides/causal_discovery.html#api): instantiate, fit, and read the result. Switching algorithms means changing only the class. Pass what you already know about the domain as [required or forbidden edges](https://pgmpy.org/guides/causal_discovery.html#expert-knowledge) with `ExpertKnowledge`. For queries, Variable Elimination is [the default choice](https://pgmpy.org/guides/probabilistic_inference.html#exact-inference) while the model is small enough for exact inference. + +Feature-engine is for work where [pandas and scikit-learn are your main tools](https://feature-engine.trainindata.com/en/latest/#sitting-at-the-interface-of-pandas-and-scikit-learn). Each transformer takes the columns it changes in its `variables` argument, so it [applies steps to selected groups of variables](https://scikit-learn.org/stable/related_projects.html). Fit the transformers on the training set and transform both sets, as the [quick start](https://feature-engine.trainindata.com/en/latest/quickstart/index.html) does. Put them in a scikit-learn Pipeline, and your [whole feature engineering pipeline](https://feature-engine.trainindata.com/en/latest/quickstart/index.html#feature-engine-within-scikit-learn-s-pipeline) saves as one object. + +XGBoost is built to be [efficient, flexible, and portable](https://xgboost.readthedocs.io/en/stable/). The same code runs on distributed environments, and its docs cover training on Dask, Spark, and Ray. Use its scikit-learn interface, like `XGBClassifier`, so it [works with scikit-learn's tools](https://xgboost.readthedocs.io/en/stable/python/sklearn_estimator.html) such as cross-validation. For categorical columns, pass a dataframe with the `category` dtype and set [`enable_categorical`](https://xgboost.readthedocs.io/en/stable/tutorials/categorical.html#training-with-scikit-learn-interface). Its docs suggest you tune with cross-validation, then [retrain with the best parameters and early stopping](https://xgboost.readthedocs.io/en/stable/python/sklearn_estimator.html#early-stopping). To keep a model, [save it with `save_model`](https://xgboost.readthedocs.io/en/stable/tutorials/saving_model.html), since a pickle is a memory snapshot meant only for checkpoints. + +TimesFM is a forecasting model that Google Research pretrained on a large time-series corpus. It [does well zero-shot](https://research.google/blog/a-decoder-only-foundation-model-for-time-series-forecasting/) on benchmarks from many domains, so you can forecast without training a model first. Install it with the extra for your backend, and load a checkpoint from the Hugging Face Hub. The code is Apache licensed, but [the pretrained weights carry their own license](https://github.com/google-research/timesfm), so check the model card before commercial or production use. + +Every pick but TimesFM works with scikit-learn. XGBoost, LightGBM, and CatBoost ship scikit-learn estimators, Feature-engine's transformers go in a Pipeline, and pgmpy is [scikit-learn compatible where possible](https://pgmpy.org/). Independent benchmarks find no single winner among the three boosting libraries, so compare them on your own data in the same cross-validation. + +Treat a model file you didn't make like code. scikit-learn's docs say to [never load a pickle from an untrusted source](https://scikit-learn.org/stable/model_persistence.html#security-maintainability-limitations), and point to skops.io or ONNX instead. XGBoost's [security notes](https://xgboost.readthedocs.io/en/stable/security.html#use-of-python-pickle) say the same about pickles.