Skip to content
View Obika-Franklin's full-sized avatar

Block or report Obika-Franklin

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
obika-Franklin/README.md

Welcome to My GitHub Portfolio! ✨

This repository showcases a selection of projects where I have applied data analysis and data science to solve real-world problems. My passion lies in transforming data into actionable insights and building robust, user-friendly applications.


Technical Toolkit 🛠️

My projects demonstrate proficiency in:

  • Programming Languages: Python, SQL, HTML
  • Frameworks: Pandas, Numpy, Matplotlib, Seaborn, TensorFlow, Scikit-Learn, FastAPI, Plotly Dash, Ngrok, Docker
  • Databases: MongoDB, MySQL, SQLite, ChromaDB
  • Tools: Microsoft Excel, PowerPoint, PowerBI, Jupyter Notebook, Google Colab, Kaggle Notebook, Github
  • Soft Skills: Strategic Problem-Solving, Synergistic Collaboration, Excellent Time Management, Articulate Communication

Projects 💻

Problem: Financial markets demand tools for risk assessment and informed decision-making.

Solution: Developed a comprehensive system for predicting stock volatility. This project utilizes a GARCH model for forecasting, powered by a FastAPI backend for model training and prediction. The system integrates Alpha Vantage API for real-time data, stores historical data in SQLite, and provides model persistence with joblib. A user-friendly web interface built with Plotly Dash enables interactive analysis and real-time use.

Impact: Provides valuable insights for investors, supports risk management strategies, and can be integrated into automated trading systems.


Problem: Low conversion of user registrations to exam completions in online learning platforms.

Solution: Conducted a controlled A/B test using synthetic data to evaluate the effectiveness of reminder emails on admissions exam completion rates. The analysis is presented through an interactive web application built with Plotly Dash, allowing stakeholders to visualize key metrics and understand the impact of behavioral nudges.

Impact: Delivers data-driven insights to enhance user engagement and improve conversion rates in digital education platforms.


Problem: Identifying and understanding households facing credit insecurity.

Solution: Applied unsupervised machine learning (K-Means Clustering) to the 2022 U.S. Federal Reserve's Survey of Consumer Finances (SCF) data. Used Principal Component Analysis (PCA) for dimensionality reduction to identify distinct segments of credit-constrained households. The insights are presented through a dynamic web-based dashboard developed with Plotly Dash.

Impact: Provides valuable insights for marketing, public policy, and financial services to better understand and address credit insecurity.


Problem: The critical need to predict building damage from earthquakes for proactive disaster preparedness.

Solution: Developed a predictive model to identify buildings prone to severe damage before an earthquake. This project involved collecting detailed building characteristics into a SQLite database. Overcame dataset imbalance using SMOTE and Random Undersampling techniques, and trained a tuned Random Forest classifier for improved accuracy in identifying vulnerable structures.

Impact: Enables proactive reinforcement and targeted aid distribution, significantly enhancing community resilience in earthquake-prone regions.


Problem: Navigating the complexities of stock market information can be challenging for many.

Solution: Built an intelligent stock analysis chatbot designed to make market insights accessible and understandable. This Python project fetches real-time or historical stock data using the Alpha Vantage API, analyzes trends, and calculates technical indicators (e.g., SMA, RSI). Analyzed data is stored and retrieved from a ChromaDB instance, enabling Retrieval-Augmented Generation (RAG) to provide contextually relevant answers using a Gemini language model. The chatbot handles queries for single or multiple stocks and includes an evaluation mechanism for response quality.

Impact: Empowers users, regardless of technical expertise, to gain insights into stock market trends and key indicators through natural language interaction.


This project involves a comprehensive analysis of weather data collected throughout 2014 for six diverse global cities: Beijing, Brasilia, Cape Town, Delhi, London, and Moscow. By examining key meteorological parameters such as temperature, precipitation, humidity, dew point, and wind speed, we aim to uncover distinct climatic patterns, assess thermal stability, characterize precipitation regimes, understand the relationship between humidity and dew point, analyze seasonal temperature variations, and evaluate wind speed characteristics across these different geographical locations. This comparative analysis provides insights into the varied climates experienced around the world and highlights notable differences and similarities in their weather behaviors during the year 2014.


This project aimed to identify the best two-week period for a summer holiday in 2023 based on historical weather data for London Heathrow. By analyzing various weather metrics such as temperature, precipitation, sunshine duration, and wind speed, a scoring system was developed to rank the weeks of the summer season. The goal was to leverage data analysis techniques to provide insights into favorable weather patterns and recommend optimal holiday dates.


In this project, I developed a comprehensive, Excel-based methodology to model and analyze the Inflow Performance Relationship (IPR) for oil wells across different flow regimes. Specifically, I calculated and graphed the IPR for both linear (undersaturated) reservoirs (which follow Darcy's law) and non-linear (saturated/two-phase) reservoirs using industry-standard empirical correlations. For the linear regime, I demonstrated the impact of key variables, such as the skin factor, production time, and reservoir depletion, on the IPR curve. For the non-linear IPR, I carried out a comparative analysis between the Vogel and Fetkovich IPR equations, highlighting their differences in predicting the well's maximum potential (Absolute Open Flow or AOF) and determining which model is more reliable for real-world application using field test points.


In this project, I performed a Decline Curve Analysis (DCA) on historical production data from an onshore field to predict future production rates and estimate the field's economic lifespan. This empirical method is one of the most widely used techniques in petroleum engineering for estimating recoverable reserves. The goal was to provide a critical forecast of the field's remaining economic life and total recoverable reserves, which is essential for guiding investment and production planning.


In this project, I performed a comparative analysis of global socio-economic indicators from 2013. The analysis utilized data retrieved from the World Bank's development indicators API, focusing on a set of key metrics including total Gross Domestic Product (GDP), GDP per capita, total population, life expectancy, the Gini index, and the unemployment rate. Before the analysis, I conducted data cleaning and preprocessing, which involved handling missing values, filtering out non-country aggregate entries, and performing unit conversions and feature engineering (specifically computing GDP per capita). The primary goal of the study was to shed light on the complex factors that contribute to a nation's overall well-being by examining the interrelationships between these various indicators.


In this project, I conducted an in-depth analysis of international trade data using Python’s pandas library to uncover trade patterns, identify key partners, and explore seasonal trends. The data, sourced from the United Nations Comtrade Database, focused on two case studies: Nigeria’s cocoa trade in 2023 (commodity codes 1801 and 1802) and the United Kingdom’s milk trade in 2014 (commodity code 0402). Through data cleaning, transformation, aggregation, and pivot table analysis, I examined trade balances, major trading partners, and temporal dynamics to derive clear and actionable insights from these commodity markets.


This project successfully demonstrated the fundamental process of converting laboratory Pressure-Volume-Temperature (PVT) data, specifically from the differential liberation experiment, into field-ready parameters, addressing Exercise 2.2 from L. P. Dake’s Fundamentals of Reservoir Engineering textbook. The analysis commenced by identifying the optimum surface separation conditions from surface separator tests, which was critical for establishing the necessary correction factors, the Shrinkage Factor ($c_{bf}$) and the Initial Solution Gas-Oil Ratio ($R_{sif}$). Utilizing these factors, the primary field PVT parameters, Oil Formation Volume Factor ($B_o$), Solution Gas-Oil Ratio ($R_s$), and Gas Formation Volume Factor ($B_g$), were calculated as a function of pressure, providing the essential input required for accurate reservoir simulations and hydrocarbon initially in place (HCIIP) calculations. Furthermore, an alternative calculation methodology was applied, converting the conventional laboratory differential parameters ($B_{od}$ and $R_{sd}$) to the field-corrected values, thereby confirming the consistency of the results and ensuring reliable data for production forecasting.


In this project, I explored advanced data analysis techniques to scrape data from the web and extract meaningful insights from real-world datasets. I focused on providing a detailed analysis of two distinct yet equally fascinating datasets: human population projections and the financial performance of top publicly traded companies. In the first section, I explored Human Population Projections, scraping and processing data from Wikipedia to investigate how urban populations were expected to change over the coming decades. I visualized the projected growth of major global cities, highlighting which urban centers were anticipated to become the most populous by the year 2100. In the second section, I analyzed the Top Publicly Traded Companies, gathering financial data such as revenue and market capitalization for the world's leading corporations. I identified the top performers in each category and examined the relationship between a company's revenue and market valuation through visualizations and statistical measures. This analysis provided insights into demographic trends and corporate financial landscapes.


This study proposes a transferable, physics-guided machine learning framework that integrates multi-well operational parameters with advanced predictive models to improve gas production forecasting. Using publicly available data from the Volve field in the North Sea, key operational variables—including wellhead pressure, wellhead temperature, choke size, and differential pressure—were combined with feature engineering techniques such as lagging, differencing, and logarithmic transformation to capture both temporal dependencies and underlying physical behavior. Two models, Random Forest (RF) and Long Short-Term Memory (LSTM), were developed and optimized using Bayesian hyperparameter tuning with time-series cross-validation. To evaluate generalization capability, the models were trained on multiple wells and tested on a completely unseen well. Results show that while the untuned LSTM model performed poorly, hyperparameter tuning significantly improved its accuracy; however, it remained less stable under unseen conditions. In contrast, the Random Forest model demonstrated consistently strong performance across training, validation, and test datasets, achieving superior robustness and reliability. The findings indicate that gas production is predominantly influenced by pressure–temperature dynamics rather than purely temporal patterns, highlighting the effectiveness of feature-driven models over sequence-based approaches in this context. Overall, the proposed framework provides a robust and interpretable solution for multi-well gas production forecasting, with significant implications for operational decision-making and field development planning.


This project is a multi-agent climate analysis system built on Heathrow weather data. I made use of a sequential pipeline where a data agent prepares the dataset, a trend agent analyses patterns, an anomaly agent detects extreme events, and a report agent generates a structured AI summary. The system is deployed on Streamlit and uses Azure OpenAI for natural language reasoning.


Feel free to explore the individual project repositories for more in-depth details, code, and demonstrations. I am always eager to learn and contribute to impactful projects!

Let's connect! 🤝

LinkedIn

Popular repositories Loading

  1. StockAnalysisAPI StockAnalysisAPI Public

    Jupyter Notebook

  2. obika-Franklin obika-Franklin Public

  3. HypothesisTesting HypothesisTesting Public

    Jupyter Notebook

  4. StockAnalysisChatbot StockAnalysisChatbot Public

    Jupyter Notebook

  5. CreditInsecurityAnalysis CreditInsecurityAnalysis Public

    Jupyter Notebook

  6. BuildingDamageAnalysis BuildingDamageAnalysis Public

    Jupyter Notebook