Big Data Event
6 min read

Five big ideas to learn at Big Data Tech Warsaw 2020

Hello again in 2020. It’s a new year and the new, 6th edition of Big Data Tech Warsaw is coming soon! Save the date: 27th of February. We have put great effort into gathering our A-team of Big Data experts from top-tier global corporations and open-source companies, willing to share some of their vast knowledge on the latest achievements and new trends in the Big Data industry. Take a look at the highlights of the most interesting pieces we’ve arranged for you this year.

Big Data Technology Warsaw Summit

1. Ways to make large-scale ML actually work

This will be one of the hottest topics on our agenda. For example, Josh Baer will talk about their winding road to better ML infrastructure at Spotify, to make the lives of internal ML practitioners easier and more productive. As you might expect, even companies that have been using ML in their products for many years and have cutting edge ML capabilities, are continuously figuring out how to scale and operate these systems and the associated teams of software engineers involved.

This topic will be also discussed by several guest experts during our main conference panel, as well as roundtable discussions.

2. Building large-scale (real-time) data analytics platforms

Many ML models require fast access to data (e.g. to detect where the nearest taxi driver is or provide personalized product recommendations while browsing a website). This is where real-time data analytics platforms come in.

Reza Shiftehfa will share his thoughts on creating a Big Data platform at Uber to handle hundreds of petabytes with real-time access. The good news is that this platform was built with a mix of open-source technology (e.g. Hadoop, Spark, Hive, Presto, Kafka) as well as tech developed internally at Uber and later on open-sourced such as Hudi and Marmaray. This means that you can also use similar technology and techniques at your own company.

We’ll also host Yuan Jiang, who will describe interactive analytics at Alibaba. Yuan will mainly focus on their large-scale real-time data warehouse. This solution is based on Apache Flink (open-source) and is adopted internally by Search, Recommendation, and Ads products.

It’s worth noting that Apache Flink is also utilized by Humn.AI, a company that offers innovative car insurance calculated by real-time algorithms. Wojciech Indyk will present how their system is built, what the advantages and limitations of Apache Flink are, and how it can be used for use-cases such as detection of a car trip (in real-time), that might look trivial at glance, but expose some traps.

3. Using data and ML to build personalized products

Building advanced ML and real-time data platforms is only one side of the coin. The second one is actually the ability to use this data in a meaningful way.

Disney+ is a brand new streaming service (launched in November 2019) with an impressive subscriber growth. Some of you might not yet know that a team in Warsaw is working on its recommendation system. Grzegorz Puchawski, who is Head of Data Science and Recommendation at Disney Streaming Services, will share problems and lessons learned, that his team have dealt with whilst working on the recommender system, which provides personalized recommendations for ESPN+ and Disney+.

Personalization will be also covered by Tomasz Burzyński and Mateusz Krawczyk who will talk about how they personalize user experience for millions of their customers, using over 20 contact channels at Orange.

4. Migrating from on-premise to the public cloud

While many companies continue building and expanding their on-premise data platforms, we have started seeing more and more companies building hybrid platforms or moving over fully to the Cloud.

You will hear about the exciting journey from our own on-premise to the cloud done together by Truecaller and GetInData. Juliana Araujo, Fouad Alsayadi and Tomasz Żukowski will share with us some exciting tech choices they made, in order to build a robust architecture, lower costs and make their data scientists happier by migrating to the Google Cloud Platform (in a series of a few steps) using a mix of on-prem, hybrid and native cloud technology. Their scale is 150+M active users that generate 30B events a day.

For those who have been using the public cloud for a while, the presentation from Adam Kurowski and Kamil Szkoda (both StepStone) can be extremely useful. They will talk about best DevOps practices in the AWS cloud and will focus on three topics: Distributed data processing, costs optimization and security.

5. Organizing and discovering data in large data lakes

Everything that we are going to talk about during the conference wouldn’t be possible without … data. You can find many analogies between data and oil or gold. Similarly to oil and gold extraction, you also need to have efficient tools to find data. This will be the topic of a joint presentation given by ING and GetInData. Verdan Mahmood and Marek Wiewiórka will talk about how they are building an enterprise-grade data discovery and data lineage at ING. Thanks to this, ING’s data scientists can easily discover available datasets in their large data lake and trust them, thanks to powerful features such as data lineage, data quality and data profiling.

6. BONUS

In this article we’ve only mentioned about 9 presentations, while the agenda includes…33! This means that there are a lot of other useful topics that you will learn about by attending the conference (February 27th, 2020) and you can find them here.

Still not convinced? Watch the video relation from previous edition:


This article was jointly written by Adam Kawa and Mikołaj Wiśniewski.
big data
analytics
conference
Warsaw
technology
bigdatatech
bigdatatechwarsaw
getindata
machine learning
13 February 2020

Want more? Check our articles

howdoweapplyknowledgeobszar roboczy 1 4

How do we apply knowledge sharing in our teams? GetInData Guilds

Do you remember our blog post about our internal initiatives such as Lunch & Learn and internal training? If yes, that’s great! If you didn’t get the…

Read more
airbyte column selectionobszar roboczy 1 4
Tutorial

Less data, less problems: Airbyte’s column selection is finally here

The Airbyte 0.50 release has brought some exciting changes to the platform: checkpointing (so that you don’t have to start from scratch in case of…

Read more
saleslstronaobszar roboczy 1 100
Tutorial

Power of Big Data: Sales

In the first part of the series "Power of Big Data", I wrote about how Big Data can influence the development of marketing activities and how it can…

Read more
getindator green santa watching a dashboard on laptop with real 9bc272ff 58b5 400a a10d 1b1639be8b3e
Tutorial

Nailing e-commerce: all data in near real-time analytics with Snowflake Dynamic Tables & Snowflake Alerts

Black Friday, the pre-Christmas period, Valentine’s Day, Mother’s Day, Easter - all these events may be the prime time for the e-commerce and retail…

Read more
transfer legacy pipeline modern using gitlab cicd
Tutorial

How we helped our client to transfer legacy pipeline to modern one using GitLab's CI/CD - Part 3

Please dive in the third part of a blog series based on a project delivered for one of our clients. Please click part I, part II to read the…

Read more
getindator stream of data showing real time analytics in busine 68956ccf d535 47c5 aa87 1b0106a634dc
Tech News

The Evolution of Real-Time Data Streaming in Business

This blog post is based on a webinar:”Real-Time Data to Drive Business Growth and Innovation in 2024” that was held by CTO Krzysztof Zarzycki at…

Read more

Contact us

Interested in our solutions?
Contact us!

Together, we will select the best Big Data solutions for your organization and build a project that will have a real impact on your organization.


What did you find most impressive about GetInData?

They did a very good job in finding people that fitted in Acast both technically as well as culturally.
Type the form or send a e-mail: hello@getindata.com
The administrator of your personal data is GetInData Poland Sp. z o.o. with its registered seat in Warsaw (02-508), 39/20 Pulawska St. Your data is processed for the purpose of provision of electronic services in accordance with the Terms & Conditions. For more information on personal data processing and your rights please see Privacy Policy.

By submitting this form, you agree to our Terms & Conditions and Privacy Policy