09/09/2026
๐ฆ๐ค๐ ๐๐จ๐ฆ๐๐ก๐๐ฆ๐ฆ ๐๐๐๐๐๐๐ก๐๐ #๐ฌ๐ฎ ๐ผ
Letโs make this one a little harder.
Imagine you have an employees table with:
* employee_id
* employee_name
* department
* salary
Management does not want the highest-paid employee.
They want the second-highest salary in each department.
The catch?
Two employees can have the same salary.
So before you start coding, think carefully:
Do you want the second employee after sorting, or the second distinct salary?
Those are not always the same thing.
For example, if salaries in Engineering are:
120000, 120000, 110000, 95000
then the second-highest distinct salary is:
110000
not 120000.
Your challenge
Write ONE SQL query that returns the employee or employees earning the second-highest distinct salary in each department.
Bonus
Return:
1. Department
2. Employee Name
3. Salary
4. Salary Rank within Department
Rules
โ Donโt hard-code department names
โ Donโt write separate queries for each department
โ Donโt simply use MAX(salary)
โ
Your query should handle salary ties correctly
โ
It should continue working as the table grows
๐ค AI era rule:
Try solving it yourself first.
Then give your solution to AI and ask:
โDoes this correctly handle duplicate salaries and ties?โ
That is a much better use of AI than asking it to solve the challenge before you think.
Use AI as your reviewer, not your replacement.
๐ฌ Drop your SQL solution in the comments.
Iโll share the solution and explain the logic after people have had time to try it.
๐ More practical SQL, Python, Data Analytics, Machine Learning, and AI resources:
EverydayDataScience.com
Follow AI & Data With Ibrahim
Learn โข Build โข Share โข Inspire
09/09/2026
๐ฆ๐ค๐ ๐๐จ๐ฆ๐๐ก๐๐ฆ๐ฆ ๐๐๐๐๐๐๐ก๐๐ #๐ฌ๐ญ ๐ฅ
Letโs see how you think about SQLโnot just how much syntax you remember.
Imagine a healthcare organization wants to identify the highest-cost patient in each age group.
Sounds simple.
But each patient can have multiple claims.
So this:
MAX(claim_amount)
does not necessarily answer the question.
It finds the largest individual claim.
Management wants something different:
Which patient has the highest combined claim amount within each age group?
That means you need to think about the problem in stages:
โ What should be aggregated first?
โ What should you group by?
โ How will you compare patients within the same age group?
โ How will you handle ties?
๐ฏ ๐ฌ๐ข๐จ๐ฅ ๐๐๐๐๐๐๐ก๐๐
Write ONE SQL query that returns the patient with the highest total claim amount in each age group.
โญ ๐๐ข๐ก๐จ๐ฆ
Return:
1. Age Group
2. Patient ID
3. Total Claim Amount
4. Rank within Age Group
๐ฅ๐จ๐๐๐ฆ:
โ Donโt hard-code patient IDs
โ Donโt write separate queries for each age group
โ Donโt simply use MAX(claim_amount)
โ
Your query should continue working as new records are added
๐ค ๐ข๐ก๐ ๐ ๐ข๐ฅ๐ ๐ฅ๐จ๐๐โฆ
Try it yourself before asking AI.
Then give your solution to AI and ask it to critique your logic, identify edge cases, and suggest improvements.
Use AI as your reviewer, not your replacement.
Because in a real SQL interview, writing syntax is only part of the challenge.
You also need to understand what the business is actually asking.
๐ฌ Drop your SQL solution in the comments.
Iโll share the solution and explain the logic after people have had a chance to try it.
๐ Save this challenge and send it to someone learning SQL.
๐ More practical SQL, Python, Data Analytics, Machine Learning, and AI resources:
EverydayDataScience.com
Follow AI & Data With Ibrahim
Learn โข Build โข Share โข Inspire
09/08/2026
๐๐ ๐ฑ๐ผ๐ฒ๐๐ปโ๐ ๐ณ๐ถ๐
๐ฏ๐ฎ๐ฑ ๐ฑ๐ฎ๐๐ฎ.
It can actually help you make bad decisions faster.
We talk a lot about becoming AI-first.
But before adding AI to a data workflow, I think there is a more important question:
Can we trust the data feeding it?
Because if your data contains:
โ Missing values handled incorrectly
โ Duplicate records
โ Inconsistent definitions
โ Broken joins
โ Unclear ownership
โ Poor data lineage
โ Conflicting business metrics
AI does not magically make those problems disappear.
It can scale them.
Think about it this way:
โ Bad Data ร AI = Faster Bad Decisions
โ
Trusted Data ร AI = Faster Better Decisions
That is why strong AI systems still need strong data foundations.
๐๐ฒ๐ณ๐ผ๐ฟ๐ฒ ๐๐ฐ๐ฎ๐น๐ถ๐ป๐ด ๐๐ถ๐๐ต ๐๐, ๐ด๐ฒ๐ ๐๐ต๐ฒ ๐ฏ๐ฎ๐๐ถ๐ฐ๐ ๐ฟ๐ถ๐ด๐ต๐:
1๏ธโฃ Validated data
Are the values complete, accurate, and reasonable?
2๏ธโฃ Consistent definitions
Does โactive customer,โ โrevenue,โ or โhigh riskโ mean the same thing across teams?
3๏ธโฃ Clear ownership
Who is responsible when the data is wrong?
4๏ธโฃ Documented lineage
Where did the data come from, and what happened to it along the way?
5๏ธโฃ Reliable pipelines
Can we trust the process that moves and transforms the data?
6๏ธโฃ Trusted metrics
Are people making decisions from the same definitions?
7๏ธโฃ Continuous quality checks
Because yesterdayโs clean dataset can become tomorrowโs broken pipeline.
๐ค ๐๐ ๐ฐ๐ฎ๐ป ๐ต๐ฒ๐น๐ฝ.
It can:
โข Detect anomalies
โข Generate SQL
โข Clean and transform data
โข Automate quality checks
โข Build pipelines
โข Summarize results
But someone still has to decide:
โข What does โcorrectโ mean?
โข Which source should we trust?
โข Is this missing value actually a problem?
โข Is this metric defined correctly?
โข Is the output reasonable?
โข What decision should follow?
That is why I donโt think AI makes data fundamentals less important.
It makes them more important.
The faster we automate, the more expensive bad assumptions can become.
๐๐ ๐ด๐ถ๐๐ฒ๐ ๐๐ผ๐ ๐๐ฐ๐ฎ๐น๐ฒ.
๐ฌ๐ผ๐๐ฟ ๐ฑ๐ฎ๐๐ฎ ๐ฑ๐ฒ๐๐ฒ๐ฟ๐บ๐ถ๐ป๐ฒ๐ ๐๐ต๐ฎ๐ ๐๐ผ๐ ๐ฎ๐ฟ๐ฒ ๐๐ฐ๐ฎ๐น๐ถ๐ป๐ด.
๐ฌ What do you think organizations should fix first before calling themselves โAI-firstโ?
๐ Save this for your next data or AI project.
๐ More practical tutorials and cheat sheets:
EverydayDataScience.com
Follow AI & Data With Ibrahim for practical lessons on SQL, Python, Data Analytics, Data Engineering, Machine Learning, and AI.
Learn โข Build โข Share โข Inspire
09/07/2026
๐ ๐ข๐ก๐๐๐ฌ ๐๐๐ง๐ ๐๐จ๐ก๐๐๐ ๐๐ก๐ง๐๐๐ฆ
Before you analyze healthcare data, ask one simple question:
๐ช๐ต๐ฎ๐ ๐ฑ๐ผ๐ฒ๐ ๐ข๐ก๐ ๐ฟ๐ผ๐ ๐ฟ๐ฒ๐ฝ๐ฟ๐ฒ๐๐ฒ๐ป๐?
This sounds basic.
But misunderstanding it can quietly ruin an entire analysis.
Imagine you have a Patients table:
1 row = 1 patient
Patient 101 has an annual healthcare cost of $24,000.
Now you have a Claims table:
1 row = 1 claim
Patient 101 has 5 claims.
You join the two tables.
Now Patient 101 appears five times.
If you carelessly sum the patientโs annual healthcare cost after that join:
$24,000 ร 5 = $120,000 โ
The query ran successfully.
The join worked.
There was no error message.
๐๐๐ ๐๐ต๐ฒ ๐ฎ๐ป๐ฎ๐น๐๐๐ถ๐ ๐ถ๐ ๐๐ฟ๐ผ๐ป๐ด.
Why?
Because the grain changed.
Before the join:
1 row = 1 patient
After the join:
1 row = 1 claim
This becomes especially important in healthcare data because you may be working with:
โข Patient-level data
โข Encounters
โข Claims
โข Diagnoses
โข Procedures
โข Medications
And these tables do not necessarily represent data at the same level.
Before joining or aggregating healthcare data, I like to ask:
1๏ธโฃ What does one row represent?
2๏ธโฃ What uniquely identifies that row?
3๏ธโฃ Is this relationship one-to-one, one-to-many, or many-to-many?
4๏ธโฃ At what grain should my metric actually be calculated?
5๏ธโฃ Did my row count change unexpectedly after the join?
๐ค ๐๐ป๐ฑ ๐๐ต๐ถ๐ ๐บ๐ฎ๐๐๐ฒ๐ฟ๐ ๐ฒ๐๐ฒ๐ป ๐บ๐ผ๐ฟ๐ฒ ๐ถ๐ป ๐๐ต๐ฒ ๐ฎ๐ด๐ฒ ๐ผ๐ณ ๐๐.
AI can write the SQL.
AI can generate the Pandas merge.
AI can calculate the aggregation.
But if you donโt understand the grain of your data, AI can help you produce a perfectly executed wrong answer.
๐ง๐ต๐ฒ ๐ฝ๐ฟ๐ผ๐ฏ๐น๐ฒ๐บ ๐๐ฎ๐๐ปโ๐ ๐๐ต๐ฒ ๐ฆ๐ค๐.
๐ง๐ต๐ฒ ๐ฝ๐ฟ๐ผ๐ฏ๐น๐ฒ๐บ ๐๐ฎ๐ ๐๐ต๐ฒ ๐ด๐ฟ๐ฎ๐ถ๐ป.
Before asking what the numbers say, understand what one row means.
๐ฌ Have you ever had a join unexpectedly multiply your rows?
๐ Save this for your next SQL, healthcare, or data-engineering project.
๐ More practical tutorials and cheat sheets:
EverydayDataScience.com
Follow AI & Data With Ibrahim for practical lessons on SQL, Python, Data Analytics, Data Engineering, Machine Learning, and AI.
Learn โข Build โข Share โข Inspire
09/04/2026
๐ ๐ก๐จ๐๐, ๐ฎ ๐๐ฒ๐ฟ๐ผ, ๐ฎ๐ป๐ฑ ๐ฎ ๐ฏ๐น๐ฎ๐ป๐ธ ๐ฎ๐ฟ๐ฒ ๐ป๐ผ๐ ๐๐ต๐ฒ ๐๐ฎ๐บ๐ฒ ๐๐ต๐ถ๐ป๐ด.
But treating them as if they are can quietly change your analysis.
Imagine a discount column:
Customer A โ 10%
Customer B โ 0%
Customer C โ NULL
Those last two values mean very different things.
0% can mean:
We know the customer received no discount.
NULL can mean:
We donโt know the discount.
Maybe it wasnโt recorded.
Maybe it wasnโt available.
Maybe the field didnโt apply.
And then there is a blank string ('').
That might look empty to you, but depending on the system, it can still be an actual stored value rather than NULL.
Why does this matter?
Because your calculations can change.
In SQL:
COUNT(*)
counts rows.
But:
COUNT(discount)
ignores NULL values.
And this can be even more dangerous:
COALESCE(discount, 0)
It is perfectly valid SQL.
But by replacing every NULL with zero, you are making an assumption:
โMissing means no discount.โ
Is that actually true?
Maybe.
Maybe not.
The same issue appears in Python/Pandas when we use things like:
fillna(0)
The code is easy.
๐ง๐ต๐ฒ ๐ฑ๐ฒ๐ฐ๐ถ๐๐ถ๐ผ๐ป ๐ถ๐ ๐๐ต๐ฒ ๐ต๐ฎ๐ฟ๐ฑ ๐ฝ๐ฎ๐ฟ๐.
Before replacing missing values, ask:
โข Why is this value missing?
โข Does zero have a real business meaning?
โข Is the missingness itself informative?
โข Will replacement distort averages or distributions?
โข Should this be imputed, excluded, flagged, or left missing?
๐ค And this matters even more when using AI.
AI can instantly write:
df['discount'].fillna(0)
or:
COALESCE(discount, 0)
But AI needs the business context to know whether that is actually the right decision.
Missing data is not just a coding problem.
It is an interpretation problem.
๐ก๐จ๐๐ โ ๐ฌ โ ๐๐น๐ฎ๐ป๐ธ.
Understanding that small distinction can prevent some very big analytical mistakes.
๐ฌ When you encounter NULLs, what do you check before deciding how to handle them?
๐ Save this for your next SQL or Python project.
๐ More practical tutorials and cheat sheets:
EverydayDataScience.com
Follow AI & Data With Ibrahim for practical lessons on SQL, Python, Data Analytics, Data Engineering, Machine Learning, and AI.
Learn โข Build โข Share โข Inspire
08/30/2026
๐๐ ๐ต๐ฎ๐ ๐บ๐ฎ๐ฑ๐ฒ ๐ถ๐ ๐ฒ๐ฎ๐๐ถ๐ฒ๐ฟ ๐๐ต๐ฎ๐ป ๐ฒ๐๐ฒ๐ฟ ๐๐ผ ๐๐ฟ๐ถ๐๐ฒ ๐ฐ๐ผ๐ฑ๐ฒ.
Paste an error.
Describe what you want.
Seconds later, you have Python or SQL that looks ready to run.
But there is a skill becoming even more important:
Debugging.
Because AI-generated code can:
โข Run without errors but produce the wrong result
โข Reference the wrong column
โข Use an inappropriate join
โข Introduce data leakage
โข Handle NULLs incorrectly
โข Make assumptions you never specified
โข Use outdated methods or libraries
So when AI gives me code, I donโt only ask:
โDoes it run?โ
I also ask:
1๏ธโฃ Do I understand what every important step is doing?
2๏ธโฃ What assumptions did AI make?
3๏ธโฃ Are the inputs and data types what the code expects?
4๏ธโฃ Can I test the logic on a few records manually?
5๏ธโฃ Do the row counts and distributions still make sense?
6๏ธโฃ Does the output answer the actual business question?
This is why I donโt think learning Python, SQL, or statistics has become less important because of AI.
I think the opposite has happened.
๐ค AI reduces the value of simply memorizing syntax.
But it increases the value of:
โข Problem solving
โข Debugging
โข Validation
โข Critical thinking
โข Business understanding
๐ฌ๐ผ๐ ๐ฑ๐ผ๐ปโ๐ ๐ป๐ฒ๐ฒ๐ฑ ๐๐ผ ๐ฐ๐ผ๐บ๐ฝ๐ฒ๐๐ฒ ๐๐ถ๐๐ต ๐๐ ๐ฎ๐ ๐๐ฟ๐ถ๐๐ถ๐ป๐ด ๐ฐ๐ผ๐ฑ๐ฒ.
You need to become good at knowing when the code is wrong.
๐ฌ Has AI made you better at debuggingโor more dependent on generated code?
๐ Save this for your next AI-assisted project.
๐ More practical tutorials and cheat sheets:
EverydayDataScience.com
Follow AI & Data With Ibrahim for practical lessons on Python, SQL, Data Analytics, Machine Learning, Data Engineering, and AI.
Learn โข Build โข Share โข Inspire
08/27/2026
๐ ๐บ๐ฎ๐ฐ๐ต๐ถ๐ป๐ฒ ๐น๐ฒ๐ฎ๐ฟ๐ป๐ถ๐ป๐ด ๐บ๐ผ๐ฑ๐ฒ๐น ๐ฐ๐ฎ๐ป ๐ฝ๐ฒ๐ฟ๐ณ๐ผ๐ฟ๐บ ๐ด๐ฟ๐ฒ๐ฎ๐ ๐ถ๐ป ๐๐ต๐ฒ ๐ป๐ผ๐๐ฒ๐ฏ๐ผ๐ผ๐ธ ๐ฎ๐ป๐ฑ ๐๐๐ถ๐น๐น ๐ณ๐ฎ๐ถ๐น ๐ถ๐ป ๐๐ต๐ฒ ๐ฟ๐ฒ๐ฎ๐น ๐๐ผ๐ฟ๐น๐ฑ.
That is one of the biggest differences between building a model and building a reliable machine-learning system.
You can have:
โข Strong validation scores
โข Clean training data
โข Good feature engineering
โข A well-tuned model
And still run into problems after deployment.
Why?
Because production changes the game.
Here are 5 reasons a good model can fail after launch:
1๏ธโฃ ๐๐ฎ๐๐ฎ ๐ฑ๐ฟ๐ถ๐ณ๐
The data your model sees today may not look like the data it was trained on.
Customer behavior changes.
Markets change.
Products change.
Processes change.
2๏ธโฃ ๐๐ผ๐ป๐ฐ๐ฒ๐ฝ๐ ๐ฑ๐ฟ๐ถ๐ณ๐
The relationship between your inputs and target can change over time.
A pattern that predicted churn six months ago may not work the same way today.
3๏ธโฃ ๐๐ฎ๐ฑ ๐ผ๐ฟ ๐บ๐ถ๐๐๐ถ๐ป๐ด ๐ถ๐ป๐ฝ๐๐๐
APIs fail.
Columns change.
Values go missing.
Categories appear that the model has never seen before.
4๏ธโฃ ๐๐ฒ๐ฒ๐ฑ๐ฏ๐ฎ๐ฐ๐ธ ๐น๐ผ๐ผ๐ฝ๐
Your model can influence the behavior it is trying to predict.
For example, if a recommendation system keeps showing the same kind of content, future user behavior may partly reflect the modelโs own decisions.
5๏ธโฃ ๐ก๐ผ ๐บ๐ผ๐ป๐ถ๐๐ผ๐ฟ๐ถ๐ป๐ด
If nobody is watching the model after deployment, performance can quietly degrade for weeks or months.
That is why production ML needs more than:
model.fit()
and
model.predict()
You also need to monitor:
โข Input distributions
โข Prediction distributions
โข Performance metrics
โข Missing values
โข Latency
โข Errors
โข Business outcomes
๐ค ๐ช๐ต๐ฎ๐ ๐ฎ๐ฏ๐ผ๐๐ ๐๐?
AI can help you write the pipeline, generate monitoring code, detect anomalies, and summarize performance.
But someone still has to decide:
โข What should be monitored?
โข When is a change significant?
โข When should the model be retrained?
โข When should we fall back to a simpler rule?
โข Is the model still helping the business?
๐ง๐ฟ๐ฎ๐ถ๐ป๐ถ๐ป๐ด ๐ฎ ๐บ๐ผ๐ฑ๐ฒ๐น ๐ถ๐ ๐ป๐ผ๐ ๐๐ต๐ฒ ๐ณ๐ถ๐ป๐ถ๐๐ต ๐น๐ถ๐ป๐ฒ.
It is the beginning of the modelโs real test.
๐ฌ Which production ML challenge do you think is hardest: drift, monitoring, feedback loops, or data quality?
๐ Save this for your next machine-learning project.
๐ More practical tutorials and cheat sheets:
EverydayDataScience.com
Follow AI & Data With Ibrahim for practical lessons on Python, SQL, Data Analytics, Machine Learning, Data Engineering, and AI.
Learn โข Build โข Share โข Inspire
Everyday Data Science: Applied AI, Agentic Systems & AI in Africa
Practical AI, machine learning, and data science for people who build, with a focus on agentic systems and applied AI in Africa. Written by practitioners.
08/22/2026
๐ข๐ป๐ฒ ๐ฆ๐ค๐ ๐ฑ๐ฒ๐ฐ๐ถ๐๐ถ๐ผ๐ป ๐ฐ๐ฎ๐ป ๐พ๐๐ถ๐ฒ๐๐น๐ ๐ฟ๐ฒ๐บ๐ผ๐๐ฒ ๐ต๐๐ป๐ฑ๐ฟ๐ฒ๐ฑ๐ ๐ผ๐ณ ๐ฟ๐ผ๐๐ ๐ณ๐ฟ๐ผ๐บ ๐๐ผ๐๐ฟ ๐ฎ๐ป๐ฎ๐น๐๐๐ถ๐.
And sometimes, you will not even notice.
Imagine you have:
1,000 customers
and you join them to an orders table.
If you use:
INNER JOIN
you only keep customers who have matching orders.
So if 180 customers have never placed an orderโฆ
they disappear.
Your result now has 820 customers.
That may be exactly what you want.
Or it may completely change the business question.
This is why, during exploratory analysis, I often prefer to start with a:
LEFT JOIN
It keeps every customer from the left tableโeven when there is no matching order.
Then I can ask:
โข Who has never ordered?
โข Which records failed to match?
โข Are the join keys clean?
โข Are we losing important rows?
โข Is โno matchโ itself an insight?
๐ง๐ต๐ฒ ๐ฑ๐ถ๐ณ๐ณ๐ฒ๐ฟ๐ฒ๐ป๐ฐ๐ฒ ๐ถ๐ ๐๐ถ๐บ๐ฝ๐น๐ฒ:
INNER JOIN โ Keep only matches.
LEFT JOIN โ Keep everything from the left + whatever matches on the right.
Neither is automatically better.
The correct join depends on the question you are answering.
Before accepting the result of any join, I like to check:
1๏ธโฃ Row count before the join
2๏ธโฃ Row count after the join
3๏ธโฃ Number of unmatched records
4๏ธโฃ Whether the join created duplicates
5๏ธโฃ Whether the join key is actually unique
๐ค ๐๐ป๐ฑ ๐๐ต๐ฎ๐ ๐ฎ๐ฏ๐ผ๐๐ ๐๐?
AI can write your SQL join in seconds.
But if you simply say:
โJoin these two tablesโ
it may produce perfectly valid SQL that answers the wrong business question.
You still need to understand:
What should happen to records that do not match?
That is not a syntax question.
That is an analysis question.
๐๐ผ๐ผ๐ฑ ๐ฆ๐ค๐ ๐ถ๐ ๐ป๐ผ๐ ๐ท๐๐๐ ๐ฎ๐ฏ๐ผ๐๐ ๐ด๐ฒ๐๐๐ถ๐ป๐ด ๐๐ต๐ฒ ๐พ๐๐ฒ๐ฟ๐ ๐๐ผ ๐ฟ๐๐ป.
๐๐ ๐ถ๐ ๐ฎ๐ฏ๐ผ๐๐ ๐ธ๐ป๐ผ๐๐ถ๐ป๐ด ๐๐ต๐ถ๐ฐ๐ต ๐ฟ๐ผ๐๐ ๐๐ผ๐ ๐ฎ๐ฟ๐ฒ ๐ธ๐ฒ๐ฒ๐ฝ๐ถ๐ป๐ด โ ๐ฎ๐ป๐ฑ ๐๐ต๐ถ๐ฐ๐ต ๐ผ๐ป๐ฒ๐ ๐๐ผ๐ ๐ฎ๐ฟ๐ฒ ๐น๐ผ๐๐ถ๐ป๐ด.
๐ฌ Which do you use more during analysis: INNER JOIN or LEFT JOIN?
๐ Save this for your next SQL project.
๐ More practical tutorials and cheat sheets:
EverydayDataScience.com
Follow AI & Data With Ibrahim for practical lessons on SQL, Python, Data Analytics, Data Engineering, Machine Learning, and AI.
Learn โข Build โข Share โข Inspire
08/21/2026
๐ ๐ผ๐ฟ๐ฒ ๐ฐ๐ต๐ฎ๐ฟ๐๐ ๐ฑ๐ผ ๐ป๐ผ๐ ๐บ๐ฎ๐ธ๐ฒ ๐ฎ ๐ฏ๐ฒ๐๐๐ฒ๐ฟ ๐ฑ๐ฎ๐๐ต๐ฏ๐ผ๐ฎ๐ฟ๐ฑ.
In fact, they often make it worse.
One of the easiest mistakes to make in Power BI, Tableau, or Excel is trying to show everything.
Every KPI.
Every chart.
Every slicer.
Every possible breakdown.
The result?
A dashboard that looks busy but makes the decision harder.
๐ ๐ด๐ผ๐ผ๐ฑ ๐ฑ๐ฎ๐๐ต๐ฏ๐ผ๐ฎ๐ฟ๐ฑ ๐๐ต๐ผ๐๐น๐ฑ ๐ฎ๐ป๐๐๐ฒ๐ฟ ๐ฎ ๐ณ๐ฒ๐ ๐ถ๐บ๐ฝ๐ผ๐ฟ๐๐ฎ๐ป๐ ๐พ๐๐ฒ๐๐๐ถ๐ผ๐ป๐ ๐พ๐๐ถ๐ฐ๐ธ๐น๐.
Before adding another visual, ask:
1๏ธโฃ ๐ช๐ต๐ฎ๐ ๐ฑ๐ฒ๐ฐ๐ถ๐๐ถ๐ผ๐ป ๐ฑ๐ผ๐ฒ๐ ๐๐ต๐ถ๐ ๐ฐ๐ต๐ฎ๐ฟ๐ ๐๐๐ฝ๐ฝ๐ผ๐ฟ๐?
If the answer is โnone,โ it probably does not belong there.
2๏ธโฃ ๐๐ผ๐ฒ๐ ๐ถ๐ ๐ฎ๐ฑ๐ฑ ๐ป๐ฒ๐ ๐ถ๐ป๐ณ๐ผ๐ฟ๐บ๐ฎ๐๐ถ๐ผ๐ป?
Three visuals telling the same story are usually unnecessary.
3๏ธโฃ ๐๐ฎ๐ป ๐๐ต๐ฒ ๐๐๐ฒ๐ฟ ๐๐ป๐ฑ๐ฒ๐ฟ๐๐๐ฎ๐ป๐ฑ ๐ถ๐ ๐ถ๐ป ๐ฑ ๐๐ฒ๐ฐ๐ผ๐ป๐ฑ๐?
If not, simplify it.
4๏ธโฃ ๐๐ ๐๐ต๐ฒ ๐บ๐ผ๐๐ ๐ถ๐บ๐ฝ๐ผ๐ฟ๐๐ฎ๐ป๐ ๐ถ๐ป๐๐ถ๐ด๐ต๐ ๐ผ๐ฏ๐๐ถ๐ผ๐๐?
Your audience should not have to hunt for the point.
5๏ธโฃ ๐๐ผ๐๐น๐ฑ ๐ผ๐ป๐ฒ ๐ฏ๐ฒ๐๐๐ฒ๐ฟ ๐๐ถ๐๐๐ฎ๐น ๐ฟ๐ฒ๐ฝ๐น๐ฎ๐ฐ๐ฒ ๐๐๐ผ ๐ผ๐ฟ ๐๐ต๐ฟ๐ฒ๐ฒ ๐๐ฒ๐ฎ๐ธ๐ฒ๐ฟ ๐ผ๐ป๐ฒ๐?
Often, yes.
๐ค ๐๐ ๐ฐ๐ฎ๐ป ๐ฏ๐๐ถ๐น๐ฑ ๐๐ถ๐๐๐ฎ๐น๐ ๐ณ๐ฎ๐๐.
But speed can make it easier to overbuild.
AI can suggest charts, layouts, and KPIs.
You still need to decide what actually deserves attention.
๐ง๐ต๐ฒ ๐ฏ๐ฒ๐๐ ๐ฑ๐ฎ๐๐ต๐ฏ๐ผ๐ฎ๐ฟ๐ฑ ๐ถ๐ ๐ป๐ผ๐ ๐๐ต๐ฒ ๐ผ๐ป๐ฒ ๐๐ถ๐๐ต ๐๐ต๐ฒ ๐บ๐ผ๐๐ ๐ฐ๐ต๐ฎ๐ฟ๐๐.
It is the one that helps someone understand the situation and make a decision faster.
๐ฌ What is the most common dashboard mistake you see?
๐ Save this before building your next dashboard.
๐ More practical tutorials and cheat sheets:
EverydayDataScience.com
Follow AI & Data With Ibrahim for practical lessons on Power BI, Excel, SQL, Python, Data Analytics, Machine Learning, and AI.
Learn โข Build โข Share โข Inspire
08/15/2026
๐ **AI can clean your data.**
But can you tell if it cleaned it correctly?
That's becoming one of the most valuable skills in data analytics.
Today, tools like ChatGPT, Claude, and GitHub Copilot can generate Pandas code in seconds.
I use AI every day.
But here's the reality:
**AI is only as good as the person reviewing its output.**
Imagine AI gives you code to fill missing values.
Do you know:
- Should the missing values be removed or filled?
- Is the median a better choice than the mean?
- Are those duplicates actually duplicates?
- Is that outlier an error or a legitimate business event?
- Did the cleaning introduce bias into your analysis?
These are decisions AI **cannot** make for you without context.
That's why every Data Analyst needs strong data cleaning fundamentals.
Master these 10 tasks and you'll be able to:
โ
Handle missing values correctly
โ
Remove duplicate records
โ
Fix incorrect data types
โ
Standardize inconsistent text
โ
Detect outliers
โ
Validate your data
Because here's the truth:
**Garbage In = Garbage Out.**
Even the most advanced AI model can't produce reliable insights from poor-quality data.
Clean data is the foundation of accurate dashboards, trustworthy reports, and successful machine learning models.
The goal isn't to avoid AI.
The goal is to **use AI with understanding.**
When you know the fundamentals, AI becomes your assistantโnot your replacement.
๐ฌ **Which data cleaning task do you find the most challenging?**
๐ **Save this post.** It's a checklist you'll use throughout your data analytics journey.
๐ Get more free Data Analytics cheat sheets, tutorials, and career resources:
**https://everydaydatascience.com**
Follow **AI & Data With Ibrahim** for practical lessons on SQL, Python, Data Analytics, Data Engineering, Machine Learning, and AI.
**Learn โข Build โข Share โข Inspire**