Description du cours
Intitulé de l'Unité d'Enseignement
Data Science et IA
Code de l'Unité d'Enseignement
21MQ061
Année académique
2026 - 2027
Cycle
MASTER
Nombre de crédits
5
Nombre heures
60
Quadrimestre
1
Pondération
Site
Montgomery
Langue d'enseignement
Français
Enseignant responsable
DENDONCKER Valentin
Objectifs et contribution de l'Unité d'Enseignement au programme
L’unité d’enseignement est une introduction aux techniques quantitatives d’exploration et d’interprétations des données, ainsi qu'aux techniques de préparation des données, en vue de les utiliser dans le cadre d'un projet impliquant des algorithmes de machine learning ou d'intelligence artificielle.
À l’issue du cours, l’étudiant sera à même de choisir et d’appliquer une technique quantitative lui permettant de répondre à une question posée à partir de données existantes.
Prérequis et corequis
Description du contenu
Le cours abordera les thèmes suivants :
- Data preparation techniques
- Supervised and unsupervised machine learning techniques
Méthodes pédagogiques
Exposés ex cathedra mêlant la description des fondements théoriques, l'illustration et l'implémentation des notions abordées.
L'intelligence artificielle peut être utilisée à des fins pédagogiques, en vue d'accompagner l'apprentissage de façon plus personnalisée.
Mode d'évaluation
L’examen écrit (à livres fermés) est composé de QCM-QRM, ainsi que d’éventuelles questions ouvertes.
L'usage de l'IA n'est pas autorisé durant l'examen.
Références bibliographiques
Andridge, R. R., & Little, R. J. A. (2010). A review of hot deck imputation for survey non-response. International Statistical Review, 78(1), 40–64.
Breiman, L. (2001). Random forests. Machine Learning, 45(1), 5–32.
Breiman, L., Friedman, J. H., Olshen, R. A., & Stone, C. J. (1984). Classification and regression trees. Wadsworth.
Breunig, M. M., Kriegel, H.-P., Ng, R. T., & Sander, J. (2000). LOF: Identifying density-based local outliers. In Proceedings of the 2000 ACM SIGMOD International Conference on Management of Data (pp. 93–104). ACM.
Brier, G. W. (1950). Verification of forecasts expressed in terms of probability. Monthly Weather Review, 78(1), 1–3.
Cawley, G. C., & Talbot, N. L. C. (2010). On over-fitting in model selection and subsequent selection bias in performance evaluation. Journal of Machine Learning Research, 11, 2079–2107.
Chandola, V., Banerjee, A., & Kumar, V. (2009). Anomaly detection: A survey. ACM Computing Surveys, 41(3), Article 15.
Cover, T. M., & Thomas, J. A. (2006). Elements of information theory (2nd ed.). Wiley.
Ester, M., Kriegel, H.-P., Sander, J., & Xu, X. (1996). A density-based algorithm for discovering clusters in large spatial databases with noise. In Proceedings of the Second International Conference on Knowledge Discovery and Data Mining (pp. 226–231). AAAI Press.
Fawcett, T. (2006). An introduction to ROC analysis. Pattern Recognition Letters, 27(8), 861–874.
Friedman, J. H. (2001). Greedy function approximation: A gradient boosting machine. The Annals of Statistics, 29(5), 1189–1232.
Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep learning. MIT Press.
Grinsztajn, L., Oyallon, E., & Varoquaux, G. (2022). Why do tree-based models still outperform deep learning on typical tabular data? Advances in Neural Information Processing Systems, 35, 507–520.
Guyon, I., & Elisseeff, A. (2003). An introduction to variable and feature selection. Journal of Machine Learning Research, 3, 1157–1182.
Hastie, T., Tibshirani, R., & Friedman, J. (2009). The elements of statistical learning: Data mining, inference, and prediction (2nd ed.). Springer.
Hernán, M. A., & Robins, J. M. (2020). Causal inference: What if. Chapman & Hall/CRC.
James, G., Witten, D., Hastie, T., Tibshirani, R., & Taylor, J. (2023). An introduction to statistical learning: With applications in Python. Springer.
Jolliffe, I. T., & Cadima, J. (2016). Principal component analysis: A review and recent developments. Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences, 374(2065), Article 20150202.
Kuhn, M., & Johnson, K. (2019). Feature engineering and selection: A practical approach for predictive models. CRC Press.
Little, R. J. A., & Rubin, D. B. (2019). Statistical analysis with missing data (3rd ed.). Wiley.
Liu, F. T., Ting, K. M., & Zhou, Z.-H. (2008). Isolation forest. In Proceedings of the 2008 Eighth IEEE International Conference on Data Mining (pp. 413–422). IEEE.
Lloyd, S. P. (1982). Least squares quantization in PCM. IEEE Transactions on Information Theory, 28(2), 129–137.
McKinney, W. (2022). Python for data analysis: Data wrangling with pandas, NumPy, and Jupyter (3rd ed.). O’Reilly Media.
Osborne, J. W. (2013). Best practices in data cleaning: A complete guide to everything you need to do before and after collecting your data. SAGE Publications.
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., & Duchesnay, É. (2011). Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12, 2825–2830.
Ross, B. C. (2014). Mutual information between discrete and continuous data sets. PLOS ONE, 9(2), e87357.
Rousseeuw, P. J. (1987). Silhouettes: A graphical aid to the interpretation and validation of cluster analysis. Journal of Computational and Applied Mathematics, 20, 53–65.
Rumelhart, D. E., Hinton, G. E., & Williams, R. J. (1986). Learning representations by back-propagating errors. Nature, 323(6088), 533–536.
Sculley, D., Holt, G., Golovin, D., Davydov, E., Phillips, T., Ebner, D., Chaudhary, V., Young, M., Crespo, J.-F., & Dennison, D. (2015). Hidden technical debt in machine learning systems. Advances in Neural Information Processing Systems, 28, 2503–2511.
van Buuren, S. (2018). Flexible imputation of missing data (2nd ed.). Chapman & Hall/CRC.
International Organization for Standardization. (2006). Environmental management — Life cycle assessment — Principles and framework (ISO 14040:2006).
National Institute of Standards and Technology. (2023). Artificial intelligence risk management framework (AI RMF 1.0) (NIST AI 100-1).
World Resources Institute, & World Business Council for Sustainable Development. (2011). Product life cycle accounting and reporting standard
pandas. (s. d.). pandas documentation.
scikit-learn. (s. d.). User guide.
seaborn. (s. d.). seaborn: Statistical data visualization.