Data Science + GenAI
Data Analytics + GenAI
Programs from ₹5,000
New Batch Starts: Coming Soon
Data Science

15 Data Science & Data Analytics Project Ideas for Freshers (with Free Datasets) — 2026

DataTeach.ai
October 9, 2026
Data Science Projects
Data Analytics Projects
Portfolio
Freshers
Machine Learning
Power BI
Generative AI
DataTeach.ai
15 Data Science & Data Analytics Project Ideas for Freshers (with Free Datasets) — 2026

Recruiters see hundreds of resumes with the same certificates. What makes yours stand out is projects — proof that you can take messy data, find answers and explain them clearly. Here are 15 project ideas for freshers, from beginner to advanced, all using free public datasets.

How to pick the right projects

  • Aiming for Data Analyst roles? Focus on dashboards, EDA and SQL-based analysis (projects 3, 4, 5, 10, 11).
  • Aiming for Data Scientist roles? Add machine learning projects (1, 2, 6, 7, 8, 13).
  • Interested in Generative AI? Build at least one LLM project (15).

Three to five well-explained projects are better than fifteen half-finished ones.

15 project ideas with datasets

1. Titanic Survival Prediction

Dataset: Kaggle – Titanic: Machine Learning from Disaster
Level: Beginner
What you'll learn: Classification basics: data cleaning, feature engineering, logistic regression and decision trees.

2. House Price Prediction

Dataset: Kaggle – House Prices: Advanced Regression Techniques
Level: Beginner
What you'll learn: Regression, handling missing values and categorical features, model evaluation with RMSE.

3. Sales Performance Dashboard

Dataset: Sample Superstore dataset (Tableau sample data, widely available on Kaggle)
Level: Beginner
What you'll learn: Excel/Power BI dashboards, KPIs, regional and category analysis.

4. Customer Churn Analysis

Dataset: Kaggle – Telco Customer Churn (IBM sample)
Level: Intermediate
What you'll learn: EDA, churn drivers, classification, and business recommendations.

5. HR Attrition Analysis

Dataset: Kaggle – IBM HR Analytics Employee Attrition & Performance
Level: Intermediate
What you'll learn: Exploratory analysis, Power BI storytelling, attrition prediction.

6. Online Retail Customer Segmentation

Dataset: UCI Machine Learning Repository – Online Retail
Level: Intermediate
What you'll learn: RFM analysis and K-Means clustering for marketing segments.

7. Movie Recommendation System

Dataset: GroupLens – MovieLens
Level: Intermediate
What you'll learn: Collaborative filtering and content-based recommendations.

8. Credit Card Fraud Detection

Dataset: Kaggle – Credit Card Fraud Detection
Level: Intermediate
What you'll learn: Imbalanced classification, precision/recall, ROC and PR curves.

9. Movie Review Sentiment Analysis

Dataset: IMDB Dataset of 50K Movie Reviews (Kaggle)
Level: Intermediate
What you'll learn: Text preprocessing, TF-IDF, classification and an introduction to NLP.

10. Netflix Content Analysis

Dataset: Kaggle – Netflix Movies and TV Shows
Level: Beginner
What you'll learn: EDA and visualisation of content trends by genre, country and year.

11. Restaurant Analysis

Dataset: Kaggle – Zomato Bangalore Restaurants
Level: Beginner
What you'll learn: EDA on ratings, cuisines, costs and locations — relatable Indian data.

12. Air Quality Analysis for Indian Cities

Dataset: Kaggle – Air Quality Data in India
Level: Intermediate
What you'll learn: Time-series EDA, seasonality and city comparisons.

13. Retail Sales Forecasting

Dataset: Kaggle – Walmart Recruiting: Store Sales Forecasting
Level: Advanced
What you'll learn: Time-series forecasting, feature engineering with holidays and seasonality.

14. Diabetes Prediction

Dataset: Kaggle – Pima Indians Diabetes Database
Level: Beginner
What you'll learn: Classification, feature scaling and model comparison.

15. Chat with Your PDFs (RAG App)

Dataset: Your own documents (college notes, policy PDFs)
Level: Advanced
What you'll learn: Generative AI: embeddings, vector database, retrieval-augmented generation and a simple chat UI.

How to present your projects to recruiters

  1. Start with the business question — “Which customers are likely to churn and why?”
  2. Show your process — data cleaning, EDA, approach and evaluation.
  3. Share results in numbers — “The model identifies 78% of churners” (use your real results).
  4. Give recommendations — what should the business do?
  5. Publish it — GitHub with a clear README, a Power BI/Tableau public link, and a short LinkedIn post.

Common mistakes

Build projects with mentor guidance

At DataTeach.ai, every program includes real-world projects reviewed by trainers. Explore Data Analytics + Generative AI, Data Science + Generative AI or Generative AI — or join a free live demo class where we build a mini project live.