FC Barcelona Sierra Leone

FC Barcelona Sierra Leone

Share

๐Ÿ‡ธ๐Ÿ‡ฑ Data Scientist | AI Educator | Founder, RiseAfrica Foundation for STEM & Innovation. Learn โ€ข Build โ€ข Share โ€ข Inspire.

Helping Africans learn AI, Data Science & technology to solve real-world problems.

09/09/2026

๐—ฆ๐—ค๐—Ÿ ๐—•๐—จ๐—ฆ๐—œ๐—ก๐—˜๐—ฆ๐—ฆ ๐—–๐—›๐—”๐—Ÿ๐—Ÿ๐—˜๐—ก๐—š๐—˜ #๐Ÿฌ๐Ÿฎ ๐Ÿ’ผ

Letโ€™s make this one a little harder.

Imagine you have an employees table with:

* employee_id
* employee_name
* department
* salary

Management does not want the highest-paid employee.

They want the second-highest salary in each department.

The catch?

Two employees can have the same salary.

So before you start coding, think carefully:

Do you want the second employee after sorting, or the second distinct salary?

Those are not always the same thing.

For example, if salaries in Engineering are:

120000, 120000, 110000, 95000

then the second-highest distinct salary is:

110000

not 120000.

Your challenge

Write ONE SQL query that returns the employee or employees earning the second-highest distinct salary in each department.

Bonus

Return:

1. Department
2. Employee Name
3. Salary
4. Salary Rank within Department

Rules

โŒ Donโ€™t hard-code department names
โŒ Donโ€™t write separate queries for each department
โŒ Donโ€™t simply use MAX(salary)
โœ… Your query should handle salary ties correctly
โœ… It should continue working as the table grows

๐Ÿค– AI era rule:

Try solving it yourself first.

Then give your solution to AI and ask:

โ€œDoes this correctly handle duplicate salaries and ties?โ€

That is a much better use of AI than asking it to solve the challenge before you think.

Use AI as your reviewer, not your replacement.

๐Ÿ’ฌ Drop your SQL solution in the comments.

Iโ€™ll share the solution and explain the logic after people have had time to try it.

๐ŸŒ More practical SQL, Python, Data Analytics, Machine Learning, and AI resources:

EverydayDataScience.com

Follow AI & Data With Ibrahim

Learn โ€ข Build โ€ข Share โ€ข Inspire

09/09/2026

๐—ฆ๐—ค๐—Ÿ ๐—•๐—จ๐—ฆ๐—œ๐—ก๐—˜๐—ฆ๐—ฆ ๐—–๐—›๐—”๐—Ÿ๐—Ÿ๐—˜๐—ก๐—š๐—˜ #๐Ÿฌ๐Ÿญ ๐Ÿฅ

Letโ€™s see how you think about SQLโ€”not just how much syntax you remember.

Imagine a healthcare organization wants to identify the highest-cost patient in each age group.

Sounds simple.

But each patient can have multiple claims.

So this:

MAX(claim_amount)

does not necessarily answer the question.

It finds the largest individual claim.

Management wants something different:

Which patient has the highest combined claim amount within each age group?

That means you need to think about the problem in stages:

โ†’ What should be aggregated first?
โ†’ What should you group by?
โ†’ How will you compare patients within the same age group?
โ†’ How will you handle ties?

๐ŸŽฏ ๐—ฌ๐—ข๐—จ๐—ฅ ๐—–๐—›๐—”๐—Ÿ๐—Ÿ๐—˜๐—ก๐—š๐—˜

Write ONE SQL query that returns the patient with the highest total claim amount in each age group.

โญ ๐—•๐—ข๐—ก๐—จ๐—ฆ

Return:

1. Age Group
2. Patient ID
3. Total Claim Amount
4. Rank within Age Group

๐—ฅ๐—จ๐—Ÿ๐—˜๐—ฆ:

โŒ Donโ€™t hard-code patient IDs
โŒ Donโ€™t write separate queries for each age group
โŒ Donโ€™t simply use MAX(claim_amount)
โœ… Your query should continue working as new records are added

๐Ÿค– ๐—ข๐—ก๐—˜ ๐— ๐—ข๐—ฅ๐—˜ ๐—ฅ๐—จ๐—Ÿ๐—˜โ€ฆ

Try it yourself before asking AI.

Then give your solution to AI and ask it to critique your logic, identify edge cases, and suggest improvements.

Use AI as your reviewer, not your replacement.

Because in a real SQL interview, writing syntax is only part of the challenge.

You also need to understand what the business is actually asking.

๐Ÿ’ฌ Drop your SQL solution in the comments.

Iโ€™ll share the solution and explain the logic after people have had a chance to try it.

๐Ÿ“Œ Save this challenge and send it to someone learning SQL.

๐ŸŒ More practical SQL, Python, Data Analytics, Machine Learning, and AI resources:

EverydayDataScience.com

Follow AI & Data With Ibrahim

Learn โ€ข Build โ€ข Share โ€ข Inspire

09/08/2026

๐—”๐—œ ๐—ฑ๐—ผ๐—ฒ๐˜€๐—ปโ€™๐˜ ๐—ณ๐—ถ๐˜… ๐—ฏ๐—ฎ๐—ฑ ๐—ฑ๐—ฎ๐˜๐—ฎ.

It can actually help you make bad decisions faster.

We talk a lot about becoming AI-first.

But before adding AI to a data workflow, I think there is a more important question:

Can we trust the data feeding it?

Because if your data contains:

โ†’ Missing values handled incorrectly
โ†’ Duplicate records
โ†’ Inconsistent definitions
โ†’ Broken joins
โ†’ Unclear ownership
โ†’ Poor data lineage
โ†’ Conflicting business metrics

AI does not magically make those problems disappear.

It can scale them.

Think about it this way:

โŒ Bad Data ร— AI = Faster Bad Decisions

โœ… Trusted Data ร— AI = Faster Better Decisions

That is why strong AI systems still need strong data foundations.

๐—•๐—ฒ๐—ณ๐—ผ๐—ฟ๐—ฒ ๐˜€๐—ฐ๐—ฎ๐—น๐—ถ๐—ป๐—ด ๐˜„๐—ถ๐˜๐—ต ๐—”๐—œ, ๐—ด๐—ฒ๐˜ ๐˜๐—ต๐—ฒ ๐—ฏ๐—ฎ๐˜€๐—ถ๐—ฐ๐˜€ ๐—ฟ๐—ถ๐—ด๐—ต๐˜:

1๏ธโƒฃ Validated data
Are the values complete, accurate, and reasonable?

2๏ธโƒฃ Consistent definitions
Does โ€œactive customer,โ€ โ€œrevenue,โ€ or โ€œhigh riskโ€ mean the same thing across teams?

3๏ธโƒฃ Clear ownership
Who is responsible when the data is wrong?

4๏ธโƒฃ Documented lineage
Where did the data come from, and what happened to it along the way?

5๏ธโƒฃ Reliable pipelines
Can we trust the process that moves and transforms the data?

6๏ธโƒฃ Trusted metrics
Are people making decisions from the same definitions?

7๏ธโƒฃ Continuous quality checks
Because yesterdayโ€™s clean dataset can become tomorrowโ€™s broken pipeline.

๐Ÿค– ๐—”๐—œ ๐—ฐ๐—ฎ๐—ป ๐—ต๐—ฒ๐—น๐—ฝ.

It can:

โ€ข Detect anomalies
โ€ข Generate SQL
โ€ข Clean and transform data
โ€ข Automate quality checks
โ€ข Build pipelines
โ€ข Summarize results

But someone still has to decide:

โ€ข What does โ€œcorrectโ€ mean?
โ€ข Which source should we trust?
โ€ข Is this missing value actually a problem?
โ€ข Is this metric defined correctly?
โ€ข Is the output reasonable?
โ€ข What decision should follow?

That is why I donโ€™t think AI makes data fundamentals less important.

It makes them more important.

The faster we automate, the more expensive bad assumptions can become.

๐—”๐—œ ๐—ด๐—ถ๐˜ƒ๐—ฒ๐˜€ ๐˜†๐—ผ๐˜‚ ๐˜€๐—ฐ๐—ฎ๐—น๐—ฒ.

๐—ฌ๐—ผ๐˜‚๐—ฟ ๐—ฑ๐—ฎ๐˜๐—ฎ ๐—ฑ๐—ฒ๐˜๐—ฒ๐—ฟ๐—บ๐—ถ๐—ป๐—ฒ๐˜€ ๐˜„๐—ต๐—ฎ๐˜ ๐˜†๐—ผ๐˜‚ ๐—ฎ๐—ฟ๐—ฒ ๐˜€๐—ฐ๐—ฎ๐—น๐—ถ๐—ป๐—ด.

๐Ÿ’ฌ What do you think organizations should fix first before calling themselves โ€œAI-firstโ€?

๐Ÿ“Œ Save this for your next data or AI project.

๐ŸŒ More practical tutorials and cheat sheets:

EverydayDataScience.com

Follow AI & Data With Ibrahim for practical lessons on SQL, Python, Data Analytics, Data Engineering, Machine Learning, and AI.

Learn โ€ข Build โ€ข Share โ€ข Inspire

09/07/2026

๐— ๐—ข๐—ก๐——๐—”๐—ฌ ๐——๐—”๐—ง๐—” ๐—™๐—จ๐—ก๐——๐—”๐— ๐—˜๐—ก๐—ง๐—”๐—Ÿ๐—ฆ

Before you analyze healthcare data, ask one simple question:

๐—ช๐—ต๐—ฎ๐˜ ๐—ฑ๐—ผ๐—ฒ๐˜€ ๐—ข๐—ก๐—˜ ๐—ฟ๐—ผ๐˜„ ๐—ฟ๐—ฒ๐—ฝ๐—ฟ๐—ฒ๐˜€๐—ฒ๐—ป๐˜?

This sounds basic.

But misunderstanding it can quietly ruin an entire analysis.

Imagine you have a Patients table:

1 row = 1 patient

Patient 101 has an annual healthcare cost of $24,000.

Now you have a Claims table:

1 row = 1 claim

Patient 101 has 5 claims.

You join the two tables.

Now Patient 101 appears five times.

If you carelessly sum the patientโ€™s annual healthcare cost after that join:

$24,000 ร— 5 = $120,000 โŒ

The query ran successfully.

The join worked.

There was no error message.

๐—•๐˜‚๐˜ ๐˜๐—ต๐—ฒ ๐—ฎ๐—ป๐—ฎ๐—น๐˜†๐˜€๐—ถ๐˜€ ๐—ถ๐˜€ ๐˜„๐—ฟ๐—ผ๐—ป๐—ด.

Why?

Because the grain changed.

Before the join:

1 row = 1 patient

After the join:

1 row = 1 claim

This becomes especially important in healthcare data because you may be working with:

โ€ข Patient-level data
โ€ข Encounters
โ€ข Claims
โ€ข Diagnoses
โ€ข Procedures
โ€ข Medications

And these tables do not necessarily represent data at the same level.

Before joining or aggregating healthcare data, I like to ask:

1๏ธโƒฃ What does one row represent?

2๏ธโƒฃ What uniquely identifies that row?

3๏ธโƒฃ Is this relationship one-to-one, one-to-many, or many-to-many?

4๏ธโƒฃ At what grain should my metric actually be calculated?

5๏ธโƒฃ Did my row count change unexpectedly after the join?

๐Ÿค– ๐—”๐—ป๐—ฑ ๐˜๐—ต๐—ถ๐˜€ ๐—บ๐—ฎ๐˜๐˜๐—ฒ๐—ฟ๐˜€ ๐—ฒ๐˜ƒ๐—ฒ๐—ป ๐—บ๐—ผ๐—ฟ๐—ฒ ๐—ถ๐—ป ๐˜๐—ต๐—ฒ ๐—ฎ๐—ด๐—ฒ ๐—ผ๐—ณ ๐—”๐—œ.

AI can write the SQL.

AI can generate the Pandas merge.

AI can calculate the aggregation.

But if you donโ€™t understand the grain of your data, AI can help you produce a perfectly executed wrong answer.

๐—ง๐—ต๐—ฒ ๐—ฝ๐—ฟ๐—ผ๐—ฏ๐—น๐—ฒ๐—บ ๐˜„๐—ฎ๐˜€๐—ปโ€™๐˜ ๐˜๐—ต๐—ฒ ๐—ฆ๐—ค๐—Ÿ.

๐—ง๐—ต๐—ฒ ๐—ฝ๐—ฟ๐—ผ๐—ฏ๐—น๐—ฒ๐—บ ๐˜„๐—ฎ๐˜€ ๐˜๐—ต๐—ฒ ๐—ด๐—ฟ๐—ฎ๐—ถ๐—ป.

Before asking what the numbers say, understand what one row means.

๐Ÿ’ฌ Have you ever had a join unexpectedly multiply your rows?

๐Ÿ“Œ Save this for your next SQL, healthcare, or data-engineering project.

๐ŸŒ More practical tutorials and cheat sheets:

EverydayDataScience.com

Follow AI & Data With Ibrahim for practical lessons on SQL, Python, Data Analytics, Data Engineering, Machine Learning, and AI.

Learn โ€ข Build โ€ข Share โ€ข Inspire

09/04/2026

๐—” ๐—ก๐—จ๐—Ÿ๐—Ÿ, ๐—ฎ ๐˜‡๐—ฒ๐—ฟ๐—ผ, ๐—ฎ๐—ป๐—ฑ ๐—ฎ ๐—ฏ๐—น๐—ฎ๐—ป๐—ธ ๐—ฎ๐—ฟ๐—ฒ ๐—ป๐—ผ๐˜ ๐˜๐—ต๐—ฒ ๐˜€๐—ฎ๐—บ๐—ฒ ๐˜๐—ต๐—ถ๐—ป๐—ด.

But treating them as if they are can quietly change your analysis.

Imagine a discount column:

Customer A โ†’ 10%
Customer B โ†’ 0%
Customer C โ†’ NULL

Those last two values mean very different things.

0% can mean:

We know the customer received no discount.

NULL can mean:

We donโ€™t know the discount.

Maybe it wasnโ€™t recorded.

Maybe it wasnโ€™t available.

Maybe the field didnโ€™t apply.

And then there is a blank string ('').

That might look empty to you, but depending on the system, it can still be an actual stored value rather than NULL.

Why does this matter?

Because your calculations can change.

In SQL:

COUNT(*)

counts rows.

But:

COUNT(discount)

ignores NULL values.

And this can be even more dangerous:

COALESCE(discount, 0)

It is perfectly valid SQL.

But by replacing every NULL with zero, you are making an assumption:

โ€œMissing means no discount.โ€

Is that actually true?

Maybe.

Maybe not.

The same issue appears in Python/Pandas when we use things like:

fillna(0)

The code is easy.

๐—ง๐—ต๐—ฒ ๐—ฑ๐—ฒ๐—ฐ๐—ถ๐˜€๐—ถ๐—ผ๐—ป ๐—ถ๐˜€ ๐˜๐—ต๐—ฒ ๐—ต๐—ฎ๐—ฟ๐—ฑ ๐—ฝ๐—ฎ๐—ฟ๐˜.

Before replacing missing values, ask:

โ€ข Why is this value missing?
โ€ข Does zero have a real business meaning?
โ€ข Is the missingness itself informative?
โ€ข Will replacement distort averages or distributions?
โ€ข Should this be imputed, excluded, flagged, or left missing?

๐Ÿค– And this matters even more when using AI.

AI can instantly write:

df['discount'].fillna(0)

or:

COALESCE(discount, 0)

But AI needs the business context to know whether that is actually the right decision.

Missing data is not just a coding problem.

It is an interpretation problem.

๐—ก๐—จ๐—Ÿ๐—Ÿ โ‰  ๐Ÿฌ โ‰  ๐—•๐—น๐—ฎ๐—ป๐—ธ.

Understanding that small distinction can prevent some very big analytical mistakes.

๐Ÿ’ฌ When you encounter NULLs, what do you check before deciding how to handle them?

๐Ÿ“Œ Save this for your next SQL or Python project.

๐ŸŒ More practical tutorials and cheat sheets:

EverydayDataScience.com

Follow AI & Data With Ibrahim for practical lessons on SQL, Python, Data Analytics, Data Engineering, Machine Learning, and AI.

Learn โ€ข Build โ€ข Share โ€ข Inspire

Photos from FC Barcelona Sierra Leone's post 08/30/2026

๐—”๐—œ ๐—ต๐—ฎ๐˜€ ๐—บ๐—ฎ๐—ฑ๐—ฒ ๐—ถ๐˜ ๐—ฒ๐—ฎ๐˜€๐—ถ๐—ฒ๐—ฟ ๐˜๐—ต๐—ฎ๐—ป ๐—ฒ๐˜ƒ๐—ฒ๐—ฟ ๐˜๐—ผ ๐˜„๐—ฟ๐—ถ๐˜๐—ฒ ๐—ฐ๐—ผ๐—ฑ๐—ฒ.

Paste an error.

Describe what you want.

Seconds later, you have Python or SQL that looks ready to run.

But there is a skill becoming even more important:

Debugging.

Because AI-generated code can:

โ€ข Run without errors but produce the wrong result
โ€ข Reference the wrong column
โ€ข Use an inappropriate join
โ€ข Introduce data leakage
โ€ข Handle NULLs incorrectly
โ€ข Make assumptions you never specified
โ€ข Use outdated methods or libraries

So when AI gives me code, I donโ€™t only ask:

โ€œDoes it run?โ€

I also ask:

1๏ธโƒฃ Do I understand what every important step is doing?

2๏ธโƒฃ What assumptions did AI make?

3๏ธโƒฃ Are the inputs and data types what the code expects?

4๏ธโƒฃ Can I test the logic on a few records manually?

5๏ธโƒฃ Do the row counts and distributions still make sense?

6๏ธโƒฃ Does the output answer the actual business question?

This is why I donโ€™t think learning Python, SQL, or statistics has become less important because of AI.

I think the opposite has happened.

๐Ÿค– AI reduces the value of simply memorizing syntax.

But it increases the value of:

โ€ข Problem solving
โ€ข Debugging
โ€ข Validation
โ€ข Critical thinking
โ€ข Business understanding

๐—ฌ๐—ผ๐˜‚ ๐—ฑ๐—ผ๐—ปโ€™๐˜ ๐—ป๐—ฒ๐—ฒ๐—ฑ ๐˜๐—ผ ๐—ฐ๐—ผ๐—บ๐—ฝ๐—ฒ๐˜๐—ฒ ๐˜„๐—ถ๐˜๐—ต ๐—”๐—œ ๐—ฎ๐˜ ๐˜„๐—ฟ๐—ถ๐˜๐—ถ๐—ป๐—ด ๐—ฐ๐—ผ๐—ฑ๐—ฒ.

You need to become good at knowing when the code is wrong.

๐Ÿ’ฌ Has AI made you better at debuggingโ€”or more dependent on generated code?

๐Ÿ“Œ Save this for your next AI-assisted project.

๐ŸŒ More practical tutorials and cheat sheets:

EverydayDataScience.com

Follow AI & Data With Ibrahim for practical lessons on Python, SQL, Data Analytics, Machine Learning, Data Engineering, and AI.

Learn โ€ข Build โ€ข Share โ€ข Inspire

Everyday Data Science: Applied AI, Agentic Systems & AI in Africa 08/27/2026

๐—” ๐—บ๐—ฎ๐—ฐ๐—ต๐—ถ๐—ป๐—ฒ ๐—น๐—ฒ๐—ฎ๐—ฟ๐—ป๐—ถ๐—ป๐—ด ๐—บ๐—ผ๐—ฑ๐—ฒ๐—น ๐—ฐ๐—ฎ๐—ป ๐—ฝ๐—ฒ๐—ฟ๐—ณ๐—ผ๐—ฟ๐—บ ๐—ด๐—ฟ๐—ฒ๐—ฎ๐˜ ๐—ถ๐—ป ๐˜๐—ต๐—ฒ ๐—ป๐—ผ๐˜๐—ฒ๐—ฏ๐—ผ๐—ผ๐—ธ ๐—ฎ๐—ป๐—ฑ ๐˜€๐˜๐—ถ๐—น๐—น ๐—ณ๐—ฎ๐—ถ๐—น ๐—ถ๐—ป ๐˜๐—ต๐—ฒ ๐—ฟ๐—ฒ๐—ฎ๐—น ๐˜„๐—ผ๐—ฟ๐—น๐—ฑ.

That is one of the biggest differences between building a model and building a reliable machine-learning system.

You can have:

โ€ข Strong validation scores
โ€ข Clean training data
โ€ข Good feature engineering
โ€ข A well-tuned model

And still run into problems after deployment.

Why?

Because production changes the game.

Here are 5 reasons a good model can fail after launch:

1๏ธโƒฃ ๐——๐—ฎ๐˜๐—ฎ ๐—ฑ๐—ฟ๐—ถ๐—ณ๐˜
The data your model sees today may not look like the data it was trained on.

Customer behavior changes.

Markets change.

Products change.

Processes change.

2๏ธโƒฃ ๐—–๐—ผ๐—ป๐—ฐ๐—ฒ๐—ฝ๐˜ ๐—ฑ๐—ฟ๐—ถ๐—ณ๐˜
The relationship between your inputs and target can change over time.

A pattern that predicted churn six months ago may not work the same way today.

3๏ธโƒฃ ๐—•๐—ฎ๐—ฑ ๐—ผ๐—ฟ ๐—บ๐—ถ๐˜€๐˜€๐—ถ๐—ป๐—ด ๐—ถ๐—ป๐—ฝ๐˜‚๐˜๐˜€
APIs fail.

Columns change.

Values go missing.

Categories appear that the model has never seen before.

4๏ธโƒฃ ๐—™๐—ฒ๐—ฒ๐—ฑ๐—ฏ๐—ฎ๐—ฐ๐—ธ ๐—น๐—ผ๐—ผ๐—ฝ๐˜€
Your model can influence the behavior it is trying to predict.

For example, if a recommendation system keeps showing the same kind of content, future user behavior may partly reflect the modelโ€™s own decisions.

5๏ธโƒฃ ๐—ก๐—ผ ๐—บ๐—ผ๐—ป๐—ถ๐˜๐—ผ๐—ฟ๐—ถ๐—ป๐—ด
If nobody is watching the model after deployment, performance can quietly degrade for weeks or months.

That is why production ML needs more than:

model.fit()

and

model.predict()

You also need to monitor:

โ€ข Input distributions
โ€ข Prediction distributions
โ€ข Performance metrics
โ€ข Missing values
โ€ข Latency
โ€ข Errors
โ€ข Business outcomes

๐Ÿค– ๐—ช๐—ต๐—ฎ๐˜ ๐—ฎ๐—ฏ๐—ผ๐˜‚๐˜ ๐—”๐—œ?

AI can help you write the pipeline, generate monitoring code, detect anomalies, and summarize performance.

But someone still has to decide:

โ€ข What should be monitored?
โ€ข When is a change significant?
โ€ข When should the model be retrained?
โ€ข When should we fall back to a simpler rule?
โ€ข Is the model still helping the business?

๐—ง๐—ฟ๐—ฎ๐—ถ๐—ป๐—ถ๐—ป๐—ด ๐—ฎ ๐—บ๐—ผ๐—ฑ๐—ฒ๐—น ๐—ถ๐˜€ ๐—ป๐—ผ๐˜ ๐˜๐—ต๐—ฒ ๐—ณ๐—ถ๐—ป๐—ถ๐˜€๐—ต ๐—น๐—ถ๐—ป๐—ฒ.

It is the beginning of the modelโ€™s real test.

๐Ÿ’ฌ Which production ML challenge do you think is hardest: drift, monitoring, feedback loops, or data quality?

๐Ÿ“Œ Save this for your next machine-learning project.

๐ŸŒ More practical tutorials and cheat sheets:

EverydayDataScience.com

Follow AI & Data With Ibrahim for practical lessons on Python, SQL, Data Analytics, Machine Learning, Data Engineering, and AI.

Learn โ€ข Build โ€ข Share โ€ข Inspire

Everyday Data Science: Applied AI, Agentic Systems & AI in Africa Practical AI, machine learning, and data science for people who build, with a focus on agentic systems and applied AI in Africa. Written by practitioners.

08/22/2026

๐—ข๐—ป๐—ฒ ๐—ฆ๐—ค๐—Ÿ ๐—ฑ๐—ฒ๐—ฐ๐—ถ๐˜€๐—ถ๐—ผ๐—ป ๐—ฐ๐—ฎ๐—ป ๐—พ๐˜‚๐—ถ๐—ฒ๐˜๐—น๐˜† ๐—ฟ๐—ฒ๐—บ๐—ผ๐˜ƒ๐—ฒ ๐—ต๐˜‚๐—ป๐—ฑ๐—ฟ๐—ฒ๐—ฑ๐˜€ ๐—ผ๐—ณ ๐—ฟ๐—ผ๐˜„๐˜€ ๐—ณ๐—ฟ๐—ผ๐—บ ๐˜†๐—ผ๐˜‚๐—ฟ ๐—ฎ๐—ป๐—ฎ๐—น๐˜†๐˜€๐—ถ๐˜€.

And sometimes, you will not even notice.

Imagine you have:

1,000 customers

and you join them to an orders table.

If you use:

INNER JOIN

you only keep customers who have matching orders.

So if 180 customers have never placed an orderโ€ฆ

they disappear.

Your result now has 820 customers.

That may be exactly what you want.

Or it may completely change the business question.

This is why, during exploratory analysis, I often prefer to start with a:

LEFT JOIN

It keeps every customer from the left tableโ€”even when there is no matching order.

Then I can ask:

โ€ข Who has never ordered?
โ€ข Which records failed to match?
โ€ข Are the join keys clean?
โ€ข Are we losing important rows?
โ€ข Is โ€œno matchโ€ itself an insight?

๐—ง๐—ต๐—ฒ ๐—ฑ๐—ถ๐—ณ๐—ณ๐—ฒ๐—ฟ๐—ฒ๐—ป๐—ฐ๐—ฒ ๐—ถ๐˜€ ๐˜€๐—ถ๐—บ๐—ฝ๐—น๐—ฒ:

INNER JOIN โ†’ Keep only matches.

LEFT JOIN โ†’ Keep everything from the left + whatever matches on the right.

Neither is automatically better.

The correct join depends on the question you are answering.

Before accepting the result of any join, I like to check:

1๏ธโƒฃ Row count before the join
2๏ธโƒฃ Row count after the join
3๏ธโƒฃ Number of unmatched records
4๏ธโƒฃ Whether the join created duplicates
5๏ธโƒฃ Whether the join key is actually unique

๐Ÿค– ๐—”๐—ป๐—ฑ ๐˜„๐—ต๐—ฎ๐˜ ๐—ฎ๐—ฏ๐—ผ๐˜‚๐˜ ๐—”๐—œ?

AI can write your SQL join in seconds.

But if you simply say:

โ€œJoin these two tablesโ€

it may produce perfectly valid SQL that answers the wrong business question.

You still need to understand:

What should happen to records that do not match?

That is not a syntax question.

That is an analysis question.

๐—š๐—ผ๐—ผ๐—ฑ ๐—ฆ๐—ค๐—Ÿ ๐—ถ๐˜€ ๐—ป๐—ผ๐˜ ๐—ท๐˜‚๐˜€๐˜ ๐—ฎ๐—ฏ๐—ผ๐˜‚๐˜ ๐—ด๐—ฒ๐˜๐˜๐—ถ๐—ป๐—ด ๐˜๐—ต๐—ฒ ๐—พ๐˜‚๐—ฒ๐—ฟ๐˜† ๐˜๐—ผ ๐—ฟ๐˜‚๐—ป.

๐—œ๐˜ ๐—ถ๐˜€ ๐—ฎ๐—ฏ๐—ผ๐˜‚๐˜ ๐—ธ๐—ป๐—ผ๐˜„๐—ถ๐—ป๐—ด ๐˜„๐—ต๐—ถ๐—ฐ๐—ต ๐—ฟ๐—ผ๐˜„๐˜€ ๐˜†๐—ผ๐˜‚ ๐—ฎ๐—ฟ๐—ฒ ๐—ธ๐—ฒ๐—ฒ๐—ฝ๐—ถ๐—ป๐—ด โ€” ๐—ฎ๐—ป๐—ฑ ๐˜„๐—ต๐—ถ๐—ฐ๐—ต ๐—ผ๐—ป๐—ฒ๐˜€ ๐˜†๐—ผ๐˜‚ ๐—ฎ๐—ฟ๐—ฒ ๐—น๐—ผ๐˜€๐—ถ๐—ป๐—ด.

๐Ÿ’ฌ Which do you use more during analysis: INNER JOIN or LEFT JOIN?

๐Ÿ“Œ Save this for your next SQL project.

๐ŸŒ More practical tutorials and cheat sheets:

EverydayDataScience.com

Follow AI & Data With Ibrahim for practical lessons on SQL, Python, Data Analytics, Data Engineering, Machine Learning, and AI.

Learn โ€ข Build โ€ข Share โ€ข Inspire

08/21/2026

๐— ๐—ผ๐—ฟ๐—ฒ ๐—ฐ๐—ต๐—ฎ๐—ฟ๐˜๐˜€ ๐—ฑ๐—ผ ๐—ป๐—ผ๐˜ ๐—บ๐—ฎ๐—ธ๐—ฒ ๐—ฎ ๐—ฏ๐—ฒ๐˜๐˜๐—ฒ๐—ฟ ๐—ฑ๐—ฎ๐˜€๐—ต๐—ฏ๐—ผ๐—ฎ๐—ฟ๐—ฑ.

In fact, they often make it worse.

One of the easiest mistakes to make in Power BI, Tableau, or Excel is trying to show everything.

Every KPI.

Every chart.

Every slicer.

Every possible breakdown.

The result?

A dashboard that looks busy but makes the decision harder.

๐—” ๐—ด๐—ผ๐—ผ๐—ฑ ๐—ฑ๐—ฎ๐˜€๐—ต๐—ฏ๐—ผ๐—ฎ๐—ฟ๐—ฑ ๐˜€๐—ต๐—ผ๐˜‚๐—น๐—ฑ ๐—ฎ๐—ป๐˜€๐˜„๐—ฒ๐—ฟ ๐—ฎ ๐—ณ๐—ฒ๐˜„ ๐—ถ๐—บ๐—ฝ๐—ผ๐—ฟ๐˜๐—ฎ๐—ป๐˜ ๐—พ๐˜‚๐—ฒ๐˜€๐˜๐—ถ๐—ผ๐—ป๐˜€ ๐—พ๐˜‚๐—ถ๐—ฐ๐—ธ๐—น๐˜†.

Before adding another visual, ask:

1๏ธโƒฃ ๐—ช๐—ต๐—ฎ๐˜ ๐—ฑ๐—ฒ๐—ฐ๐—ถ๐˜€๐—ถ๐—ผ๐—ป ๐—ฑ๐—ผ๐—ฒ๐˜€ ๐˜๐—ต๐—ถ๐˜€ ๐—ฐ๐—ต๐—ฎ๐—ฟ๐˜ ๐˜€๐˜‚๐—ฝ๐—ฝ๐—ผ๐—ฟ๐˜?

If the answer is โ€œnone,โ€ it probably does not belong there.

2๏ธโƒฃ ๐——๐—ผ๐—ฒ๐˜€ ๐—ถ๐˜ ๐—ฎ๐—ฑ๐—ฑ ๐—ป๐—ฒ๐˜„ ๐—ถ๐—ป๐—ณ๐—ผ๐—ฟ๐—บ๐—ฎ๐˜๐—ถ๐—ผ๐—ป?

Three visuals telling the same story are usually unnecessary.

3๏ธโƒฃ ๐—–๐—ฎ๐—ป ๐˜๐—ต๐—ฒ ๐˜‚๐˜€๐—ฒ๐—ฟ ๐˜‚๐—ป๐—ฑ๐—ฒ๐—ฟ๐˜€๐˜๐—ฎ๐—ป๐—ฑ ๐—ถ๐˜ ๐—ถ๐—ป ๐Ÿฑ ๐˜€๐—ฒ๐—ฐ๐—ผ๐—ป๐—ฑ๐˜€?

If not, simplify it.

4๏ธโƒฃ ๐—œ๐˜€ ๐˜๐—ต๐—ฒ ๐—บ๐—ผ๐˜€๐˜ ๐—ถ๐—บ๐—ฝ๐—ผ๐—ฟ๐˜๐—ฎ๐—ป๐˜ ๐—ถ๐—ป๐˜€๐—ถ๐—ด๐—ต๐˜ ๐—ผ๐—ฏ๐˜ƒ๐—ถ๐—ผ๐˜‚๐˜€?

Your audience should not have to hunt for the point.

5๏ธโƒฃ ๐—–๐—ผ๐˜‚๐—น๐—ฑ ๐—ผ๐—ป๐—ฒ ๐—ฏ๐—ฒ๐˜๐˜๐—ฒ๐—ฟ ๐˜ƒ๐—ถ๐˜€๐˜‚๐—ฎ๐—น ๐—ฟ๐—ฒ๐—ฝ๐—น๐—ฎ๐—ฐ๐—ฒ ๐˜๐˜„๐—ผ ๐—ผ๐—ฟ ๐˜๐—ต๐—ฟ๐—ฒ๐—ฒ ๐˜„๐—ฒ๐—ฎ๐—ธ๐—ฒ๐—ฟ ๐—ผ๐—ป๐—ฒ๐˜€?

Often, yes.

๐Ÿค– ๐—”๐—œ ๐—ฐ๐—ฎ๐—ป ๐—ฏ๐˜‚๐—ถ๐—น๐—ฑ ๐˜ƒ๐—ถ๐˜€๐˜‚๐—ฎ๐—น๐˜€ ๐—ณ๐—ฎ๐˜€๐˜.

But speed can make it easier to overbuild.

AI can suggest charts, layouts, and KPIs.

You still need to decide what actually deserves attention.

๐—ง๐—ต๐—ฒ ๐—ฏ๐—ฒ๐˜€๐˜ ๐—ฑ๐—ฎ๐˜€๐—ต๐—ฏ๐—ผ๐—ฎ๐—ฟ๐—ฑ ๐—ถ๐˜€ ๐—ป๐—ผ๐˜ ๐˜๐—ต๐—ฒ ๐—ผ๐—ป๐—ฒ ๐˜„๐—ถ๐˜๐—ต ๐˜๐—ต๐—ฒ ๐—บ๐—ผ๐˜€๐˜ ๐—ฐ๐—ต๐—ฎ๐—ฟ๐˜๐˜€.

It is the one that helps someone understand the situation and make a decision faster.

๐Ÿ’ฌ What is the most common dashboard mistake you see?

๐Ÿ“Œ Save this before building your next dashboard.

๐ŸŒ More practical tutorials and cheat sheets:

EverydayDataScience.com

Follow AI & Data With Ibrahim for practical lessons on Power BI, Excel, SQL, Python, Data Analytics, Machine Learning, and AI.

Learn โ€ข Build โ€ข Share โ€ข Inspire

08/15/2026

๐Ÿš€ **AI can clean your data.**

But can you tell if it cleaned it correctly?

That's becoming one of the most valuable skills in data analytics.

Today, tools like ChatGPT, Claude, and GitHub Copilot can generate Pandas code in seconds.

I use AI every day.

But here's the reality:

**AI is only as good as the person reviewing its output.**

Imagine AI gives you code to fill missing values.

Do you know:

- Should the missing values be removed or filled?
- Is the median a better choice than the mean?
- Are those duplicates actually duplicates?
- Is that outlier an error or a legitimate business event?
- Did the cleaning introduce bias into your analysis?

These are decisions AI **cannot** make for you without context.

That's why every Data Analyst needs strong data cleaning fundamentals.

Master these 10 tasks and you'll be able to:

โœ… Handle missing values correctly

โœ… Remove duplicate records

โœ… Fix incorrect data types

โœ… Standardize inconsistent text

โœ… Detect outliers

โœ… Validate your data

Because here's the truth:

**Garbage In = Garbage Out.**

Even the most advanced AI model can't produce reliable insights from poor-quality data.

Clean data is the foundation of accurate dashboards, trustworthy reports, and successful machine learning models.

The goal isn't to avoid AI.

The goal is to **use AI with understanding.**

When you know the fundamentals, AI becomes your assistantโ€”not your replacement.

๐Ÿ’ฌ **Which data cleaning task do you find the most challenging?**

๐Ÿ‘‡ **Save this post.** It's a checklist you'll use throughout your data analytics journey.

๐Ÿ“š Get more free Data Analytics cheat sheets, tutorials, and career resources:

**https://everydaydatascience.com**

Follow **AI & Data With Ibrahim** for practical lessons on SQL, Python, Data Analytics, Data Engineering, Machine Learning, and AI.

**Learn โ€ข Build โ€ข Share โ€ข Inspire**

Want your business to be the top-listed Gym/sports Facility in Jersey City?

Click here to claim your Sponsored Listing.

Location

Category

Website

Address


Somerset
Jersey City, NJ
08873