Am Neumarkt 😱

Machine learning and other gibberish
See also: https://sharing.leima.is
Archives: https://datumorphism.leima.is/amneumarkt/

08:52 · Jan 28, 2022 · Fri

确实有很多，比如我用 ack 替代了 grep，速度快了不少。

https://www.ruanyifeng.com/blog/2022/01/cli-alternative-tools.html

07:39 · Jan 25, 2022 · Tue

#ml

https://ruder.io/ml-highlights-2021/

ruder.io

ML and NLP Research Highlights of 2021

This post summarizes progress across multiple impactful areas in ML and NLP in 2021.

07:13 · Jan 21, 2022 · Fri

#visualization

Seaborn is getting a new interface.

Would be great if the author defines a dunder method _ _ add _ _ () instead of using .add() method. Using dunder add, we can simply use + on layers.

Nevertheless, we can all move away from plotnine when the migration is done.

https://seaborn.pydata.org/nextgen/

visualization

07:39 · Jan 20, 2022 · Thu

#ds

Deepnote supports Great Expectations (GE) now.

I ran their template notebook:

https://deepnote.com/project/Reduce-Pipeline-Debt-With-Great-Expectations-mLT9DFCQSpW4kUBAzzdhBw/%2Fnotebook.ipynb/#00000-e170fae0-7e06-4a7a-85f3-343584ec4b94

07:01 · Jan 20, 2022 · Thu

#visualization

Beautiful, elegant, and informative. It reminds me of the Netflix movie chromatic storytelling visualization.

Full image:
https://zenodo.org/record/5828349

Other discussions:
https://www.reddit.com/r/dataisbeautiful/comments/s6vh8k/dutch_astronomer_cees_bassa_took_a_photo_of_the/

visualization

21:15 · Jan 17, 2022 · Mon

#python

I thought it was a trivial talk in the beginning.
But I quickly realized that I may know every each piece of the code mentioned in the video but the philosophy is what makes it exciting.

He talked about some fundamental ideas of Python, e.g., protocols.

After watching this video, an idea came to me. Pytorch lightning has implanted a lot of hooks in a very pythonic way. This is what makes pytorch lightning easy to use. (So if you do a lot of machine learning experiments, pytorch lightning is worth a try.)

https://youtu.be/cKPlPJyQrt4

YouTube

James Powell: So you want to be a Python expert? | PyData Seattle 2017

www.pydata.org

PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each…

python

23:49 · Jan 3, 2022 · Mon

#ds

https://2022.pycon.de/blog/pyconde-pydata-berlin-tickets/

2022.pycon.de

PyConDE & PyData Berlin 2022 Tickets

Tickets for PyConDE & PyData Berlin 2022

21:26 · Dec 30, 2021 · Thu

#data #ds

Disclaimer: I'm no expert in state diagram nor statecharts.

It might be something trivial but I find this useful: Combined with some techniques in statecharts (something frontend people like a lot), state diagram is a great way to document what our data is going through in data (pre)processing.

For complicated data transformations, we can make the corresponding state diagram and follow your code to make sure it is working as expected. The only thing is that we are focusing on the state of data not any other system.

We can use some techniques from statecharts, such as hierarchies and parallels.

State diagram is better than flowchart in this scenario because we are more interested in the different states of the data. State diagrams automatically highlights the states and we can easily spot the relevant part in the diagram and we don’t have to start from the beginning.

I documented some data transformations using state diagrams already. I haven't tired but it might also help us document our ML models.

References:
1. https://statecharts.dev
2. https://en.wikipedia.org/wiki/State_diagram

statecharts.dev

Welcome to the world of Statecharts

The world of statecharts describes what statecharts are, their benefits and drawbacks, how they differ from state machines, and practical examples on how to use them.

data ds

09:26 · Dec 24, 2021 · Fri

#visualization

Pu X, Kay M. A probabilistic grammar of graphics. Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems. New York, NY, USA: ACM; 2020. doi:10.1145/3313831.3376466
Available at: https://dl.acm.org/doi/10.1145/3313831.3376466

A very good read if you are visualizing probability densities a lot.
The paper began with a common mistake people make when visualizing densities. Then they proposed a systematic grammar of graphics for probabilities. They also provide a package (quite preliminary, see here https://github.com/MUCollective/pgog ).

visualization

08:47 · Dec 18, 2021 · Sat

#ml #science

I remember several years ago when I was still doing my PhD, there's this contest about predicting protein structure and none of them was working well. At that time, I would never have thought we could have anything like AlphaFold in a few years.
.

https://www.science.org/content/article/breakthrough-2021

Science

Science’s 2021 Breakthrough of the Year: AI brings protein structures to all

Bounty of new structures will forever change biology and medicine

ml science

22:09 · Dec 14, 2021 · Tue

#visualization #fun

https://www.githubwrapped.com/

visualization fun

20:40 · Dec 14, 2021 · Tue

#ML #Transformers

Alammar J. The Illustrated Transformer. [cited 14 Dec 2021]. Available: http://jalammar.github.io/illustrated-transformer/

So good.

jalammar.github.io

The Illustrated Transformer

Discussions:
Hacker News (65 points, 4 comments), Reddit r/MachineLearning (29 points, 3 comments)

Translations: Arabic, Chinese (Simplified) 1, Chinese (Simplified) 2, French 1, French 2, Italian, Japanese, Korean, Persian, Russian, Spanish 1, Spanish…

ML Transformers

07:40 · Dec 13, 2021 · Mon

#DS #visualization

https://percival.ink/

A new lightweight language for data analysis and visualization. It looks promising.

I hate jupyter notebooks and I don't use them on most of my projects. One of the reasons is low reproducibility due to its non-reative nature. You changed some old cells and forgot to run a cell below, you may read wrong results.
This new language is reactive. If old cells are changed, related results are also updated.

percival.ink

Percival • Web-based, reactive Datalog notebooks

Percival is a declarative data query and visualization language for exploring complex datasets, producing interactive graphics, and sharing results.

DS visualization

10:19 · Dec 11, 2021 · Sat

#ml #rl

How to Train your Decision-Making AIs
https://thegradient.pub/how-to-train-your-decision-making-ais/

The author reviewed "five types of human guidance to train AIs: evaluation, preference, goals, attention, and demonstrations without action labels".

The last one reminds me of the movie Finch. In the movie, Finch was teaching the robot to walk by demonstrating walking but without "labels".

The Gradient

How to Train your Decision-Making AIs

How do humans transfer their knowledge and skills to artificial decision-making agents more efficiently? What kind of knowledge and skills should humans provide and in what format?

ml rl

09:52 · Dec 5, 2021 · Sun

#visualization

Hmmm my plate is way off the planetary heath diet recommendation.

Source:
https://www.nature.com/articles/d41586-021-03612-1

visualization

10:36 · Dec 2, 2021 · Thu

#DS

Just in case you are also struggling with Python packages on Apple M1 Macs

I am using the third option: anaconda + miniforge.

https://www.anaconda.com/blog/apple-silicon-transition

Anaconda

A Python Data Scientist’s Guide to the Apple Silicon Transition | Anaconda

Even if you are not a Mac user, you have likely heard Apple is switching from Intel CPUs to their own custom CPUs, which they refer to collectively as “Apple Silicon.” The last time Apple changed its computer architecture this dramatically was 15 years ago…

21:31 · Dec 1, 2021 · Wed

11月25号是消除对妇女的暴力行为国际日，来自metaLab的研究人员在随机选择一百万条#MeToo推文后，仔细阅读了转发次数超过 100 次的示例，在894 条推文中只有 8 条是关于性侵犯或围绕#MeToo主题的经历的实际推文，其余绝大多数是新闻媒体和政治讨论，其中大多数都忽略了#MeToo运动核心的具体问题和幸存者的声音，设计师Kim Albrecht想通过这个可视化项目来展示被忽视的针对女性暴力问题

Twitter

Kim Albrecht

Today is the International Day for the Elimination of Violence against Women. To shine some light on the enormity of the problem we visualized some unheard tweets from the #WhyIDidntReport hashtag. metoo.kimalbrecht.com #GenerationEquality #OrangeTheWorld…

21:36 · Nov 30, 2021 · Tue

#visualization

An interactive Visual Vocabulary:

https://ft-interactive.github.io/visual-vocabulary/

visualization

15:08 · Nov 29, 2021 · Mon

#tool

https://www.jetbrains.com/fleet/

JetBrains

JetBrains Fleet: The Code Editor and IDE for Any Language

Built from scratch, based on 20 years of experience developing IDEs. Fleet uses the IntelliJ code-processing engine, with a distributed IDE architecture and a reimagined UI.

tool

11:59 · Nov 19, 2021 · Fri

#ML

SHAP (SHapley Additive exPlanations) is a system of methods to interpret machine learning models.
The author of SHAP built an easy-to-use package to help us understand how the features are contributing to the machine learning model predictions. The package comes with a comprehensive tutorial for different machine learning frameworks.

- Python Package: [slundberg/shap](https://shap.readthedocs.io/)
- A tutorial on how to use it: https://www.aidancooper.co.uk/a-non-technical-guide-to-interpreting-shap-analyses/

---

The package is so popular and you might be using it already. So what is SHAP exactly? It is a series of methods based on Shapley values.

> SHAP (SHapley Additive exPlanations) is a game-theoretic approach to explain the output of any machine learning model.
>
> -- [slundberg/shap](https://github.com/slundberg/shap)

Regarding Shapley value: There are two key ideas in calculating a Shapley value.
- A method to measure the contribution to the final prediction of some certain combination of features.
- A method to combine these "contributions" into a score.

SHAP provides some methods to estimate Shapley values and also for different models.

The following two pages explain Shapley value and SHAP thoroughly.

- https://christophm.github.io/interpretable-ml-book/shap.html
- https://christophm.github.io/interpretable-ml-book/shapley.html

References:
- Lundberg SM, Lee SI. A unified approach to interpreting model predictions. of the 31st international conference on neural …. 2017. Available: http://papers.nips.cc/paper/2017/file/8a20a8621978632d76c43dfd28b67767-Paper.pdf
- Lundberg SM, Nair B, Vavilala MS, Horibe M, Eisses MJ, Adams T, et al. Explainable machine-learning predictions for the prevention of hypoxaemia during surgery. Nature Biomedical Engineering. 2018;2: 749–760. doi:10.1038/s41551-018-0304-0

---
I posted [a similar article years ago in our Chinese data weekly newsletter](https://github.com/data-com/weekly/discussions/27) but for a different story.

Aidan Cooper

Explaining Machine Learning Models: A Non-Technical Guide to Interpreting SHAP Analyses

With interpretability becoming an increasingly important requirement for machine learning projects, there's a growing need for the complex outputs of techniques such as SHAP to be communicated to non-technical stakeholders.

Before

After