TalkDEV Track
Data Science Platform with Kubeflow on Kubernetes
- When
- Length
- 25 minutes
- Track
- DEV Track
- Edition
- 10th ISTA Conference
- Where
- Free Virtual Event
- Times shown in
- Europe/Sofia
About this session
Ideally, there would be a single unified Data Science Platform, that everyone would use, for all Data Science use cases.
There are more challenges towards that ideal point. The Data Science Process has more phases each further branching out into many tasks, and each task can potentially have different hardware requirements (CPUs, GPUs, etc.), and software requirements (programming languages, libraries, frameworks, etc.). The software tools that can be used for one task are usually more, and overall the software tools landscape is very fragmented and has high variance. During development, there is constant change in data, code, and ML models. Our end-to-end Data Science Workflows need to be reliable, portable, scalable, and reusable. We should also be able to iterate fast and work Agile-like. It is not just about training ML models and prediction, there is a lot more to operationalization and lifecycle management of ML products.
The modern approach for addressing all these challenges is to use containerized components, connected in a microservices architecture, to orchestrate the entire system from one main control point, and to apply DevOps principles for continuous development, integration, deployment, and monitoring. In other words, we need MLOps.
After this talk, you will be able to start experimenting towards making your own Data Science Platform.
Published by the speaker for ISTA 2020, reproduced verbatim.
Speaker
Archive
ISTA 2020 in the archive
14 sessions were published for this edition, each with its own page, its speakers and the times as the programme printed them.
