AI Agents Full Course 2026 | AI Agents Tutorial For Beginners | Agentic AI Course | Edureka Live
Skills:
Agent Foundations90%
Key Takeaways
Builds a foundation in Agentic AI using LangChain, RAG, and LLM Ops to create autonomous AI agents
Full Transcript
Hello everyone and welcome to the AI agents full course. Your gateway to understanding how intelligent systems operate autonomously. In this course you will start by learning the core principles of AI deep learning and large language models. From there you will explore agentic AI discovering how AI agents can make decisions solve problems and execute task independently. Along the way you will gain practical knowledge of lang chain rag LLM ops and prompt engineering essential tools for building next generation AI applications. We will also highlight groundbreaking AI technologies career opportunities and strategies to prepare for interviews in this evolving fit. So before we begin, please like, share and subscribe to Edureka's YouTube channel and hit the bell icon to stay updated on the latest tech content from Edureka. Also check out Edureka's agentic AI certification training. It is carefully crafted to meet industry demands and prepare you for the future of intelligent agents. You will gain practical skills in lang LLM ops and more through live instructorled sessions and hands-on labs. Whether you're a beginner or a tech professional, this course helps you master the concepts and accelerate your AI career. So check out the course link given in the description box below. So first let us start by understanding what agentic AI is. Agentic AI is transforming industries by allowing machines to learn, adapt and evolve independently similar to live organisms. Unlike traditional AI, these intelligent agents investigate, optimize, and develop solutions over time without requiring direct human participation. Recent advancements include OpenAI's deep research, which automatically analyzes massive amounts of data to provide detailed reports, and Google's Gemini 2.0, which improves AI's capacity to plan and reason across different data types. Service Now's AI agent orchestrator is transforming enterprise automation by coordinating many AI agents to address difficult business concerns. As these systems become more powerful, they have the potential to unlock ideas beyond the human imagination ranging from wind turbine blade design to AIdriven company management. Let's start with our first topic. What is agentic AI? Agentic AI denotes artificial intelligence systems capable of autonomously executing actions to attain designated objectives unlike reactive AI which only responds to the inputs. Agentic AI is proactive capable of planning, adapting and making decisions autonomously. So let's explore deep into agentic AI and see its capabilities. Agentic AI is a type of artificial intelligence that exhibits autonomous behavior, enabling it to take actions and operate without continuous human guidance. It is goal-driven, actively working towards achieving specific objectives rather than passively responding to inputs like reactive AI. And with advanced decision-m capabilities, it can evaluate multiple options, select the optimal course of action based on current conditions and acquired knowledge and adapt its strategies dynamically in response to unforeseen changes in its environment. Moreover, agentic AI demonstrates proactiveness by taking the initiative to act rather than waiting for external triggers making it highly effective in dynamic and complex scenarios. Now let us see its relevance in the current AI market. When AI systems can act autonomously to accomplish predefined objectives, we call that agentic AI, making it highly relevant in the current AI market. Its autonomy allows it to operate without continuous human guidance, making decisions and adapting dynamically to achieve objectives. This capability is complemented by its advanced problem solving skills, enabling it to evaluate complex situations, strategize and respond effectively to challenges. However, the growing adoption of agentic AI also rises important ethical considerations such as ensuring responsible behavior, minimizing unintended consequences and maintaining transparency in its decision-m processes. Now that you know about agentic AI, so let us discuss how it differ from other AI systems. Agentic AI differs significantly from other AI systems in its autonomy, decision making and adaptability to achieve long-term goals. Unlike reactive AI which performs predefined task only when prompted such as spam filters or image classifiers, agentic AI takes the initiative and operates independently. It also contrast with the generative AI which focuses on creating content like child GPT generating text but it is not goal-driven by combining autonomous behavior, strategic decision making and the ability to adapt dynamically. Agentic AI stands out as a powerful system designed to achieve specific objectives in evolving environments. Now since we know a bit of differences, let us see the comparison between generative AI and agentic AI. Generative AI and agentic AI differ in several key aspects that define their functionality and applications. Generative AI is primarily focused on creation, excelling in output focused tasks such as generating text, images or other form of content. Its adaptability is limited as it relies heavily on prompts for guidance and lacks the ability to operate independently. In contrast, agentic AI emphasizes autonomy, making it goal-driven and capable of dynamically adapting to changing environments. Unlike the prom dependent nature of generative AI, agentic AI is self-directed, enabling it to take the initiative and execute strategic task effectively. These differences highlight the complimentary roles of both AI types in addressing distinct challenges. Now let us see the impact of agentic AI on various industries. Agentic AI has had a profound impact across various industries transforming operations and solving long-standing challenges. Autonomous logistics systems such as those in Amazon warehouses have significantly improved operational efficiency by 30 to 40%. In healthcare, AI enabled surgical robots like the Davinci system have performed over 10 million less invasive procedures worldwide, enhancing precision and patient outcomes. Scientific advancements have also been transformed by systems like Deep Minds Alpha Fold, which successfully solved the decades old protein folding problem. On a global scale, the World Economic Forum predicts that by 2025, AI will displace 85 million jobs while creating 97 million new ones, reshaping the labor market. And in the energy sector, AI powered smart grids can reduce electricity waste by up to 10%. Promoting greener energy solutions. Additionally, over 90 countries are investing in AI enabled military technology to modernize their defense systems, showcasing the strategic importance of agentic AI in global security. Now, let us see the applications of agentic AI. Agentic AI is transforming various industries by enabling systems to make autonomous decisions, adapt to changing environments, and achieve specific goals. Autonomous vehicle powers self-driving cars and drones to navigate roads, avoid obstacles, and make real-time decisions as seen with Tesla autopilot and autonomous delivery drones. In robotics, agentic AI allows industries healthcare and exploration robots to perform complex task independently as demonstrated by Boston Dynamics robots used in logistics and rescue operations. Personalized virtual assistants like Google Assistant and Amazon Alexa leverage agentic AI to predict user needs, manage schedules, and execute task without direct commands. And in gaming, adaptive AI agents enhance the experience by creating challenging humanlike opponents such as Alph Go and AI boards in the real-time strategy games. In healthcare, agentic AI supports personalized treatments, accurate diagnostics, and surgical assistance with examples including AIdriven surgical robots and systems for remote patient monitoring. These applications demonstrate the transformative potential of agentic AI across diverse domains. Agentic AI is making a significant impact across various industries by enabling autonomy, adaptability, and efficiency in diverse applications. In finance, it powers algorithmic trading systems and fraud detection tools, optimizing financial operations such as managing investment portfolios and identifying fraudulent activities. In smart cities, AI systems manage energy consumptions, optimize traffic flow and enhance public safety with examples like smart traffic lights adapting in real time and autonomous energy grid optimization. In space exploration, autonomous spacecraft and planetary rovers such as NASA's Mars rovers perform exploration task independently. In education, AI powered tutors like Carnegie Learning provide personalized instruction by adapting to individual learning styles. In military and defense, autonomous drones and surveillance system improves situational awareness and decision making such as AIdriven surveillance drones in defense applications. Now let us see the challenges and risk associated with agentic AI. While agentic AI offers tremendous potential, it also faces several challenges and risk that must be addressed to ensure its safety and ethical deployment. So one key concern is misalignment with human goals where AI system may pursue objectives that conflict with human intentions due to poorly defined parameters or intended unintended consequences such as autonomous robot prioritizing efficiency over safety. Ethical questions arise regarding accountability and decision-m demonstrated by the challenge of determining who is responsible when an autonomous vehicle causes an accident. The complexity of decision-m in agentic AI can also lead to a lack of transparency making it difficult to understand or explain its actions particularly in sensitive fields like healthcare or finance. Ensuring safety and reliability is another challenge as AI systems must operate effectively in unpredictable environments such as autonomous drones encountering extreme weather or medical failures. Additionally, agentic AI systems often require substantial computational resources making their deployment costly as seen in advanced robotics and self-driving cars. Security vulnerabilities pose further risk as autonomous systems could be targeted by cyber attacks potentially leading to harmful consequences like the manipulation of autonomous vehicles. Lastly, overdependence on AI may reduce human oversight or lead to skill degradation in critical areas such as relying too heavily on autonomous systems for medical diagnosis without human validation. These challenges highlight the need for robust design, rigorous testing and ethical frameworks to mitigate risk and maximize the benefits of agentic AI. Now let's see the future of agentic AI. The future of agentic AI is set to be transformative with advancements across various domains influencing its deployment. Future systems will exhibit increased autonomy and adaptability, enabling them to make a complex decisions in real time and operate effectively in dynamic environments without human intervention. The integration of agentic AI with advanced technologies like quantum computing, IoT, the edge computing will further enhance its capabilities allowing for faster decision making and realtime processing at the edge. These systems will have the widespread applications in sectors such as healthcare where they will enable autonomous medical diagnostics, personalized treatment plans and robotic surgery. Climate action with advanced systems for environmental monitoring and response and space exploration where smart rovers and spacecraft will carry out missions on their own. As these technologies evolve, ethical concerns and accountability will need to be addressed. promoting the development of regulatory frameworks to ensure responsive AI usage. Additionally, agentic AI will foster human AI collaboration, enhancing productivity and creativity in the fields such as education, engineering, and research. Imagine asking Chad GP for a poem and it writes one instantly. Now think about an AI assistant planning your entire day, booking meetings, and even handling emails without your constant input. That's the difference between generative AI which creates content and agentic AI which acts with autonomy making decisions. In 2025, as AI becomes more than just a tool, understanding the shift is very critical. Are we heading towards just smarter chatbots or truly independent digital agents? Let's break it down through this video. To truly understand the ship, let's first break down what generative AI is. Generative AI is a type of artificial intelligence designed to create content, whether it's text, images, music, or even code. Instead of making decisions or even taking action on its own, it focuses on producing outputs based on the patterns it has learned from the vast amounts of data. At its core, generative AI models use deep learning techniques like transformers to generate new content that resembles human created work. For example, Chad GPT generates humanlike text based on prompts. Midjenny and Dali creates stunning images from simple text description and GitHub copilots helps developers suggesting code snippets in real time. Generative AI has several strengths. It enhances creativity and productivity allowing artists, writers and programmers to work faster and even more efficient. It scales effortlessly generating unlimited variation of content in just few seconds. It also adapts responses based on user input making interactions feel more personalized. But it also comes with few limitations. Generative AI lacks autonomy. It doesn't think or act on its own. It only responds when prompted. It has no real decision-m abilities and cannot evaluate consequences or make even independent choices. Additionally, it can generate biased or inaccurate content based on the data that it has seen. While generative AI is powerful for creating, it cannot act independently. And that's where agentic AI comes in. Let's explore what agentic AI is. Agentic AI goes beyond just generating content. It acts autonomously making decisions and executing tasks without the need of constant human input. Unlike generative AI which can only responds to prompts, agentic AI can plan, adapt and take initiatives based on goals rather than the specific instructions. At its core, agentic AI combines reasoning, memory, and decision making to operate more like an independent agent. It doesn't just create, it analyzes, strategize, and acts. Real world examples include autonomous robots which navigates and complete the task on their own. AIdriven personal assistant like those managing schedules, booking flights and handling emails without human oversight. Even self-driving cars which continuously assess their environment and make split-second driving decisions. Agent AI has its own strengths. It reduces the needs for manual intervention automating the complex workflows. It adapts to real world conditions, learning and improving over time. It can even handle multi-step tasks that require planning, execution, and adjustment. But it also has its own challenges. Developing truly autonomous AI requires significant advancements in reasoning and adaptability. There are certain risks including unintended behaviors and ethical concerns around AI, which makes independent decisions. And unlike generative AI which focuses on creativity, agentic AI is limited in how well it can generate novel content. So while generative AI creates and agentic AI acts, the real powers comes when these two work together. Let's see the key differences between generative AI and agentic AI. Generative AI and agentic AI serve different purposes, each with unique strengths and applications. The key distinction comes down to creativity versus decision making. As previously discussed, generative AI focuses on producing content, whether it's text, image, or code. It enhances creativity by assisting writers, designers, and developers. But it lacks true autonomy. It only works when prompted and doesn't make any decision on its own. Agent AI, on the other hand, is designed for interactions and execution. Instead of just generating responses, it can analyze situations, make decisions, and take actions. While it may not create content like generative AI, it can manage workflows, automate task and adapt to real world conditions. Another key difference is user dependency. Generative AI is entirely reactive, meaning it requires human input to function. It waits for prompts before generating anything. In contrast, agentic AI is proactive. It can initiate actions independently, setting reminders, optimizing schedules, or even solving problems without human intervention. The applications of these AI types also differ. Generative AI is widely used in content creating, marketing, entertaining, and software development. And agentic AI powers autonomous system like self-driving cars. AI powered customer service and personal assistant that can handle complex workflows. Both AI types are transforming the industries. But when they work together, they unlock even greater potential. Imagine an AI that not only generates a marketing campaign, but also launches it, tracks engagement, and refine the strategy automatically. The future isn't just about choosing between generative AI and agentic AI. It's about combining them two to build truly intelligent systems. Now that we understand the key differences between these two, let's explore the future of AI by asking, will generative AI be replaced? As AI continues to evolve, one big question arises. Will agentic AI replace generative AI? Right now, generative AI is everywhere, helping people write, design, and code faster than ever before. But it has one major limitation. It relies entirely on human input. Agent AI on the other hand takes things further. It doesn't just generate, it decides, plans, and even acts. It's the next step towards the true autonomous intelligence. Does that means generative AI will be obsolete? Not necessarily. The future of AI isn't about one replacing the other. It's about coexisting. Generative AI will keep getting more creative and even sophisticated, producing even higher quality content. Agentic AI will become even more autonomous, integrating deeper with industries like healthcare, finance, and robotics. But this shift does comes with some risk. As AI takes on decision-making power, we face new challenges. ethical concerns, unintended consequences and the need for accountability. If an AI agent makes a bad decision, who is responsible? And how do we ensure it aligns with the human values? The answer lies in balance. The real future of AI is hybrid approach where generative AI fuels creativity and agentic AI drives intelligent action. Imagine an AI system that not only writes a research paper but also submits it to generals, responds to reviews and refine it automatically. And this is where we are headed. Not just smarter AI, but AI that truly works with us as both a creator and an agent. The question isn't whether agentic AI will replace generative AI. It's how we'll harness both to shape the future of intelligence. Now that we have explored the differences between generative AI and agenic AI, let's move on to building an intelligent AI agent that can interact with our database using natural language. This means you can simply ask a question like show me all the students who have scored about 80 and the agent will automatically convert it into an SQL query, fetch the data and return the exact result from the database. No need to write complex SQL queries manually. Just ask and the AI response. Let's dive in and build this powerful system. First, we need to set up a cond environment to manage our project dependency. To do this, we open the terminal and run the following command. We'll write create p vv python equals to 3.10 - y. So, creates a new environment and hyphen pvnv specify the environment path as VNV. Python equals to 3.10 installs Python version 3.10 inside the environment and hyphen y automatically confirms the installation without asking for approval. Once the process is complete, our virtual environment is ready and we can move forward with setting up our agentic AI project. Next, we'll create a file named requirements.txt. txt where we'll list all the necessary libraries for our project. This will help us easily install dependencies in one go. Additionally, we'll create a NV file to securely store our Google generative AI API key, keeping sensitive information separate from our main code. With these files in place, we ensure a well structured and organized setup for our agentic AI project. First, we will work with SQLite, a lightweight self-contained database engine to create and manage a student database. Let's break it down step by step. So, we'll create a file named SQL. py and import the SQLite 3 module which allows us to work with SQLite databases. We'll write import SQLite 3. This module provides all the necessary functions to create a database, insert records, retrieve data, and manage connections. Next, we create a connection to an SQLite database file named student db. We'll write connection equals to skite3 doconnect connection equals to skite3.connect in the bracket in double inverted comma student db. If this file doesn't exist, SQL light will automatically create it. The connection object will allow us to interact with the database. Now we create a cursor object which is used to execute SQL commands in Python. We'll write cursor equals to connection.cursor. Think of the cursor as a tool that helps us send queries to the database and retrieve results. Now we define a SQL command to create a table named student with four columns. We'll write table info equals to triple inverted commas. Next we'll create a table. For that we'll write create table. Then student we'll write in the bracket name type vcar and we'll have 25 characters. Comma class type vcar and the same 25 characters. Comma section type var with 25 characters and marks type integer. Then we'll write cursor.execute in the bracket table info. The name stores the students name string up to 25 characters. The class store the class's name and the section stores the section of the student and lastly the mark stores the marks obtained as integer. Executing this commands creates the table in the database. Next, we insert five student records into the student table using SQL insert statements. I've already created and inserted five values in the table. You can create as much as you can. Each insert commands adds a new role with the students name, class, section, and marks. Now, we retrieve and display all records from the student table. For that, we'll have to write print in the bracket. Print in the bracket the inserted records are. In the next line, we'll write data equals to cursor do.executed in the bracket three single inverted comma select star from student closing the inverted commas in the bracket. Then we'll write for row in data colon print in the bracket row. The select star from student query fetches all the data from the table. The for loop iterates through the records and prints them one by one. And finally we commit our changes and close the database connection. For that we'll write connection commit and then connection.c close. The dotcommit function ensures all the changes are saved in the database. The dot close closes the connection freeing up the system resources. And that's it. We have successfully created a student database inserted records and retrieved them using SQLite in Python. Now let's build an interactive stream app that converts natural language questions into SQL queries using Google's Gemini model. It then retrieves data from an SQLite database and display the result. Let's break it down step by step. But before we start, we have to activate the environment. For that, we'll write activate venv forward slash. And here our environment is activated. First, we'll create a file named app. py and load environment variables using env. For that we'll write from env we'll import load env. Next we'll write load env. It will load all environment variables. This ensures that sensitive information such as API keys is securely stored and accessed. Next we import the necessary modules. For that we'll write import streamlit as st. Then import OS. Then import escalite 3 and then import Google.generative AI as genai. Streamlight here powers the web interface. OS helps access the environment variables. SQLite 3 allows us to interact with the database and Google generative AI enables the conversion of natural language into SQL queries. Now we configure the Google Gemini API key. But before that we'll have to create a API key through Google studio itself. I've already generated one. You can create yours through Google studio itself. Then we'll write genai.configure in the bracket API_key equals to os do.get env key. This allows the app to use Gemini 1.5 Pro to generate SQL queries. Then we define a function to generate SQL queries from natural language input using gemi. For that we'll write defaf get_jemni response in the bracket question, prompt. Next we'll write model equals to genai, generative model in the bracket we'll write models/jna version 1.5 pro. Then we'll write response equals to model.generate rate underscore content in the bracket and in square brackets prompt in the square bracket zero and comma question and then we'll write return response text. The function initializes the Gemini model. It takes a question and predefined prompt as input and the AI model generates an SQL query as output. Next, we define a function to execute SQL queries on the database and retrieve results. For that we'll write def read_sql_query in the bracket sql comma db. Next we'll write con equals to skqite 3 dot connect in the bracket db. Then cur equals to con.cursor and then cur equals to execute in the bracket sql. Then we'll write rows equals to cur do fetch call. then con dot commit and then con.t close and then we'll create a loop by writing for row in rows and then we'll print it and then return rows. The function connects to the student db database. It executes the given SQL's query and it fetches all the retrieve records and prints them. Now we define the AI prompt that instructs Gemini on how to convert the questions into SQL queries. As you can see, I've already created a prompt for my own and you can create yours according to how you want your model to function. If you want the prompt which I've used over here, you can just comment on the video and I'll send it to you. This prompts ensures the Gemini AI generates SQL queries accurately without unnecessary text. Now we'll set page configuration with a title and icon. For that we'll write st set_page configuration in the bracket page title equals to SQL query generator edurea comma page icon. Then we'll display the edureka logo and header. For that we'll write st dot image in the bracket 123.png png comma width equals to let's keep it as 200 st dom markdown in the bracket logo plus ederica's gemini app/ your AI powered SQL assistant next we'll write next we'll write st.mmarkdown then the logo and ask any questions and I'll generate the SQL query for you the page title and the icon are set a logo is displayed at the top and the app's purpose is to introduce to the user. And before we import the logo, just make sure that you have the logo in your folder. We take user input for a natural language query. For that, we'll write question equals to st.ext_input in the bracket enter your query in plain English colon, key equals to input. This allows users to type their questions such as show all students with marks above 80. A submit button triggers the SQL generation process and for that we'll write submit equals to ST dot button in the bracket generate SQL query. When clicked the app processes the query and retrieves the result. Now we define what happens when the submit button is clicked. For that we'll write if submit in the next line response equals to get gemini response in the bracket question, prompt. This is to convert the question to SQL and then we'll print the response. Then we'll write response equals to read_sql_query in the bracket response, student db. And this is to execute SQL on the database. Then we'll write ST dos subheader. In the bracket the response is brackets close. Next we'll include a loop for then row in response. Then we'll write st. dot subheader in the bracket the responses and then we'll include a loop for row in response. Then we'll print row and then st dot header and in the brackets row. The user's question is converted into an SQL query using Gemini AI. The SQL query is executed on the student DB database and the retrieve records are displayed on the streamllet app and that's it. The AI powered streamllet app allows users to ask natural language questions which are automatically converted into SQL queries and executed on a student database. Now let's open the terminal and run our streamllet app. To do this, we simply type streamllet run app. py and hit enter. It's running. And as you can see, our agentic AI is up and running, ready to interact with our database. Let's test it by asking a simple question. We'll ask, give me the names of all the students. The AI processes our request, converts it into an SQL query, and retrieves the student names from the database. Perfect. As you can see, the response is generated. Now, let's try another query. We'll say, give me the average of marks. And just like that, the AI calculates and returns the average marks. The response which is provided is 72.2. So in this video, we successfully built an agentic AI that can understand natural language, generate SQL queries, and interact with our data seamlessly. Amazon just dropped a major AI upgrade, Alexa, and it's unlike anything we have seen before. It's not just an update, it's a complete transformation powered by generative AI. But what exactly makes Alexa smarter? more conversational and more capable. Well, in this video, we will break down how Amazon has leveraged state-of-the-art AI models to make Alexa a true AI assistance. How it compares to competitors like Chad GBT voice and Google Assistants, and whether it's the future of voice AI. Let's rewind a bit. Alexa started as a simple voice assistance in 2014. It could set reminders, play music, and control smart devices. But it had one major limitation. It wasn't really thinking, just following predefined rules. As AI advanced, assistants like Apple Siri and Google Assistants improved. But Amazon saw an opportunity to turn Alexa into a true conversational AI. And that's where generative AI comes in. Enter Alexa Plus, a brand new AI powered version of Alexa that understands context, remembers conversations, and sounds more natural than ever. Launched on February 26, 2025, Alexa Plus is Amazon's next generation AI assistance designed to provide more natural conversational interactions and enhanced capabilities. This upgrade enables Alexa to perform complex tasks such as planning events, managing schedules, and controlling smart home devices more efficiently. Alexa Plus represents a significant evolution from the original Alexa, introducing several key enhancements. So, let us see what are they. First, we have conversational abilities. Alexa Plus offers more natural and expansive interactions, understanding colloquial expressions and complex ideas, making conversational feel smoother and more intuitive. Building on that, it also takes a more proactive approach to assisting users. Unlike the original Alexa, which primarily responded to direct commands, Alexa Plus can anticipate user needs such as suggesting earlier dispatches due to traffic or notifying about sales on desired items. In addition, it has become more personalized than ever. Alexa Plus can remember user preferences, dietary restrictions, and important dates, tailoring responses and actions to individual needs. Whereas the original Alexa had limited personalization capabilities. Beyond personalization, it also enhances task management. The new Alexa can handle complex tasks like making reservations, ordering groceries, and coordinating multiple services seamlessly, surpassing the more basic functionalities of the original Alexa. Not just that, it also integrated with more services than before. Alexa Plus connects with a broader range of services and devices including GrubHub, Open Table, Ticket Master and various smart home products making it even more versatile. On top of all these improvements, it now has the ability to act independently. Agentic capabilities is a notable advancements in Alexa plus. Now that we have seen how Alexa plus has improved. So let's dive into the technology behind it and understand how generative AI models and agentic AI capabilities power this next generation assistance. Alexa plus is built on cuttingedge generative AI and agentic AI leveraging powerful models and algorithms to process language, understand context and execute task autonomously. So let's break down the key technologies that make this possible. Large language models which is LLM the brain behind conversations. At the core of Alexa plus is an advanced transformer-based language model similar to GPD4 cler and Amazon's preparatory Titan model. This LLMs are trained on vast data sets allowing Alexa to understand complex queries and respond naturally. also maintain context across conversations, making interactions feel more fluid and generate humanlike responses, reducing robotic and repetitive phrasing. And by using techniques like reinforcement learning with human feedback, Alexa Plus continuously improves its conversation's ability based on real world interactions. The next technology is agentic AI, enabling proactive and autonomous actions. Beyond just responding to commands, Alexa Plus integrates agentic AI models which allow it to act independently. Built on rack and action models, it can plan multi-step task, example, finding a restaurant, booking a table, and arranging transportation. It retrieves realtime web data to provide the latest information and execute action across multiple apps and services without user micromanagement. This enables a fully autonomous AI assistance experience, reducing the need for manual user input. After agentic AI, the technology that makes Alexa so versatile is neural network architectures, enabling speech and context awareness. Alexa plus utilizes deep learning techniques such as sequencetosequence models for natural language generation but which stands for birectional encoder representations from transformers for understanding user intent with greater accuracy. Next, whisper ASR automatic speech recognition for improved voice processing making Alexa more responsive to different accent and speech patterns. These advancements enable highly accurate speech recognition, contextually understanding, and real-time adaption to user behavior. Alexa Plus integrates long-term memory storage using vector databases like FISS or Amazon Aurora, allowing it to remember user preferences over time, adapt to individual habits and routines for a more personalized experience, also provide contextual reminders based on past interactions. This deep personalization is what makes Alexa plus feel more like a true digital assistant rather than just a voice control device. Then comes the technology that makes Alexa plus capable of understanding and interacting with user which is multimodel AI. Alexa plus leverages multimodel AI combining natural language processing for text based queries computer vision for eos show devices enabling it to process and analyze on screen content. Also the speech synthesis is used to generate humanlike voice responses and this makes Alexa place capable of understanding and interacting with user in multiple ways enhancing its overall functionality and by combining LLM agentic AI deep learning models and real-time data retrieval. Alexa represent a significant leap in AIdriven virtual assistance. It is no longer just a voice assistance. It is an autonomous context aware and highly personalized AI companion designed to make daily life easier. Now that we have explored the technology behind Alexa plus, so let us see how it stack up against other leading AI assistants. Alexa plus enters the AI assistant space with generative AI and agentic AI making it smarter and more proactive. But how does it compare to the top AI models available today? So let's break it down across key aspects. So we will compare them based on five key factors such as AI powered and capabilities, personalization and memory. Then proactive and autonomous task. Next is the ecosystem and third party integration and finally conversational abilities. So first let us compare it with AI power and capabilities. So how powerful is the AI behind each assistant? Alexa plus uses Amazon Titan plus custom LLM with generative AI and agentic AI for smart proactive responses. Whereas Chat GP voice runs on GP4 great for deep conversation but lacks real world task execution. Whereas Google assistants uses Gemini AI best for search and multimodel inputs such as text, voice and images. And Apple Siri uses Ajax LLM improving in language but still rule based and limited. So Alexa plus leads in proactive AI while GPT4 dominates in conversation. Next let us compare it in terms of personalization and memory. So can the assistant remember your preferences and adapt? Let us see. So Alexa plus long-term memory of routines, preferences and contextual adaption. Charge voice limited memory resets after sessions. Whereas Google assistants remembers preferences inside Google apps but lacks deep personalizations. Apple Siri is minimal memory. Mostly relies on Apple's preset commands. So, Alexa Plus leads in remembering and adapting to users. Next is the proactive and autonomous task execution. So, can it handle task on its own? Let's see. Alexa plus uses agentic AI for multi-step automation. For example, booking, ordering, and reminders. Chat GPD voice assist with planning but can't perform real world automation. Google Assistants can set reminders and retrieve information but lacks deep automation whereas Apple Siri limited to commands relies on shortcuts for basic automation. So here the Alexa Plus is the most proactive handling task automatically. Next is the ecosystem and third party integration. How well does it work with other devices and apps? Well, Alexa plus is best for smartome such as Amazon Eco Ring third party integrations. Chant deputy voice can connect to some external tools but no smart home control. Google assistance is deep integration with Google apps and services whereas Apple Siri is limited to Apple devices with minimal third party support. So here again Alexa plus and Google Assistant Elite but Alexa has better smart home control and finally conversational abilities. So how natural and humanlike are the conversations? Alexa plus is natural, expressive and context aware. Chat JD voice is best for deep intelligent conversations. Google assistance is accurate but more search focused. Apple series still command based with limited depth. So chat GPD voice is best for deep conversations but Alexa plus is most natural for voice interactions. So now let us see the future of AI assistance. Let's see what's next. So here we have smarter AI memory. assistants will remember and personalized even better. Next, more autonomy. AI will handle complex multi-step task independently. And then more human-like conversations. AI will feel more natural and intuitive. Next, seamless integration. AI will connect effortlessly across devices and services. Next, realtime decision making. AI will anticipate needs and offer proactive help. So Alexa plus is best for automation, memory and smart home control. Whereas chativity voice is best for deep intelligent conversation. Google assistance is best for search and Google productivity. Whereas Apple Siri is best for Apple users but still limited in AI features. And Alexa plus is not just an upgraded. It's a redefinition of AI assistance with generative AI and agentic AI for smarter proactive help. So what do you think which AI assistance is your favorite? Let me know in the comments below. AI is no longer just responding. It's acting, planning, and automating entire workflows. Welcome to the era of agentic AI, where AI agents can write code, run businesses, and make decisions without human input. By 2030, AI automation is projected to be a $200 billion industry. And those who master agentic AI tools like autogen AI and language chain will lead the future. Now let's dive into the ultimate road map to mastering agentic AI. So first let's see how you can build a strong foundation in generative AI. To truly master agentic AI, you need a strong foundation in generative AI. Understanding how AI models work, their evolution, and their impact on automation. So start by exploring how AI has evolved from rule-based systems to advanced models like chart GPD, autogen. You can check out Edurea's video on what is generative AI and generative AI examples for valuable insights into the fundamentals of generative AI, its real world applications and how it is transforming various industries. So first understand the core concepts of agentic AI where AI can perceive, plan and act independently to automate complex workflows. Next learn about real world applications such as business automation, AI powered software engineering and autonomous research agents. And to deepen your knowledge familiarize yourself with the key AI models like GPD4 Turbo Cloud AI, Gemini and Mistral. and stay updated on multi- aent systems and self-improving AI trends. And for hands-on exploration, leverage open AI's API, clut AI and llama 3, or experiment with different AI models on hugging face spaces. You can also stay updated with AI research papers from archive and hugging face to keep up with the latest breakthroughs. Edureka's generative AI certification and training will teach you Python programming, data science, artificial intelligence, natural language processing and so many other updated technologies that a beginner or advanced learners is seeking. And by understanding this concepts and experimenting with these tools, you will have a strong foundation to start working with agentic AI. Next, let's dive into the programming for AI. To build and experiment with agentic AI, you need to understand the fundamentals of programming, especially in Python, which is backbone of AI development. Start by learning Python basics, focusing on data structures, loops, functions, and object-oriented programming. Then explore essential AI and machine learning libraries like numpy and pandas for data manipulation. Mattplot lib and seaborn for data visualization and tensorflow and pytorch for deep learning. To work with AI agents, you must also understand API interactions as most AI tools like OpenAI's API lang chain and hugging face models require API calls. Additionally, learning automation with fast API, Flask and web scraping can help you integrate AI into real world applications. For hands-on practice, start small projects like building a chatbot, creating an AI powered summarizer or automating data analysis. You can also explore Edurea's Python training and certification course designed by industry experts where you will learn Python from scratch along with key libraries like NumPy, Pandas, Mattplot Lib and Psychit Learn through hands-on project and real world applications. With this powerful promptic techniques and tools, you will be able to optimize AI responses and unlock the full potential of agentic AI. To leverage agentic AI, start experimenting with cuttingedge tools that enable autonomous workflows. Auto GPD and Crew AI allows you to create multi- aent AI systems where AI agents collaborate to complete task. Baby AGI is perfect for automated research and decision-m helping AI iterate on task dynamically. Daving AI the first AI software engineer showcases how AI can independently write debug and deploy code for hands-on learning build real world projects like AI powered automation assistance autonomous research tools or self-improving chatbots to see agentic AI in action by working with these tools you will understand how AI can move beyond just responding to acting intelligently and autonomously. Next, explore LANC chain and rack. Powerful tools that give AI the ability to retrieve real-time information, process external data, and enhance decision making. To build more powerful and contextaware AI applications, understanding lang and rack is essential. Langchain is must-learn framework that enables seamless integration of LLM with external data sources allowing AI agents to interact with APIs, databases, and documents. Rag enhances AI models by providing memory and real-time knowledge retrieval, making responses more accurate and upto-date. For hands-on learning, try building your own AI chatbot with lang capable of retrieving real-time information instead of relying on static training data. A great project idea to explore is an AI powered research assistant capable of summarizing papers, fetching real world data, and answering domain specific questions. And to dive deeper, check out our dedicated video on lang chain and rag where we cover everything in detail. Next, here are the extra tips for your success. To excel in agentic AI, consistent practice and community engagement are key. So, start by pushing your AI projects to GitHub and using version control like Git to track your progress and collaborate. Join AI communities on Discord, Twitter, and Hugging Face spaces where you can interact with experts and stay updated on trends and get feedback on your work. Take advantages of AI internships and open-source projects to gain real world experience and build a strong portfolio. Also stay updated by regularly reading AI research papers on archive and Google Scholar, keeping up with the latest advancements in multi- aent AI and automation. And by following these extra tips, you will accelerate your AI learning with career growth. Think about this. Instead of you doing all your work, you have a machine to finish it for you or it can do something which you thought was not possible. For instance, predicting the future like predicting earthquake, tsunami so that preventive measures can be taken to save lives, chat bots, virtual personal assistants like Siri in iPhones, Google Assistant and believe me, it is getting smarter day by day with deep learning. self-driving cars. It will be a blessing for elderly people and disabled people who find it difficult to drive on their own. And on top of that, it can also avoid a lot of accidents that happen due to human error. Google AI eye doctor. So this is a recent initiative by Google where Google is working with an Indian eye care chain to develop an AI software which can examine retina scans to identify a condition called diabetic retinopathy which can cause blindness. AI music composer. Who thought that we can have an AI music composer using deep learning? And maybe in the coming years even machines will start winning Grammys. And one of my favorites, a dream reading machine with so many unrealistic applications of AI and deep learning that we have seen so far. I was wondering that whether we can capture dreams in the form of a video or something. And I wasn't surprised to find out that this was tried in Japan a few years back on three test subjects and they were able to achieve close to 60% accuracy and that is amazing but I'm not sure that whether people would want to be a test subject for this or not because it can reveal all your dreams. Great. So this sets the base for you and we are ready to understand what is artificial intelligence. Artificial intelligence is nothing but the capability of a machine to imitate intelligent human behavior. AI is achieved by mimicking a human brain by understanding how it thinks, how it learns and work while trying to solve a problem. For example, a machine playing chess or a voice activated software which helps you with various things in your phone or a number plate recognition system which captures the number plate of an oversp speeding car and processes it to extract the registration number and identify the owner of the car so that he can be charged and all of these wasn't very easy to implement before deep learning. Now let's understand the various subsets of artificial intelligence. So till now you'd have heard a lot about artificial intelligence, machine learning and deep learning. However, do you know the relationship between all three of them? So deep learning is a sub field of a sub field of artificial intelligence. So it is a sub field of machine learning which is a sub field of artificial intelligence. So when we look at something like Alph Go, it is often portrayed as a big success for deep learning. But it's actually a combination of ideas from several different areas of AI and machine learning like deep learning, reinforcement learning, self-play etc. And the idea behind deep neural networks is not new but it dates back to 1950s. However, it became possible to practically implement it only when we had the new high-end resource capability. So I hope that you have understood what is artificial intelligence. So let's explore machine learning followed by its limitations. So machine learning is a subset of artificial intelligence which provide computers with the ability to learn without being explicitly programmed. In machine learning, we do not have to define all the steps or conditions like any other programming application. However, we have to train the machine on a training data set large enough to create a model which helps the machine to take decisions based on its learning. For example, if we have to determine the species of a flower using machine, then first we need to train the machine using a flower data set which contains various characteristics of different flowers along with the respective species as you can see here in the image. We have got the sele length, sele width, petal length, petal width and the species of the flower too. So using this input data set, the machine will create a model which can be used to classify a flower. Next, we'll pass on a set of characteristics as input to the model and it will output the name of the flower. And this process of training a machine to create a model and use it for decision making is called machine learning. However, this process had some limitations. Machine learning is not capable of handling highdimensional data that is where input and output is large and it is present in multiple dimensions and handling and processing such a data becomes very complex and resource exhaustive and this is termed as the curse of dimensionality. So to understand this in simpler terms, let us consider a line of 100 yards and let us assume that you drop the coin somewhere in the line. You'll easily find the coin by simply walking on the line. A line is a single dimension entity. Now let's consider that you have got a square of side 100 yards each and you dropped a coin somewhere inside the square. Now definitely you'll take more time to find the coin within that square. A square is a two-dimensional entity. Now let's take it a step ahead and consider a cube of side 100 yard each and you dropped a coin somewhere inside the cube. Now it is even more difficult to find the coin. So if we see that the complexity is increasing as the dimensions are increasing and in real life the highdimensional data that we're talking about has got many dimensions which makes it very very complex to handle and process. The highdimensional data can be easily found in use cases like image processing, natural language processing, image translation etc. And machine learning was not capable of solving this use cases and hence deep learning came to the rescue. So deep learning is capable of handling the highdimensional data and is also efficient in focusing on the right features on its own. And this process is called feature extraction. Now let's try and understand how deep learning works. So in an attempt to re-engineer a human brain, deep learning studies the basic unit of a brain called a brain cell or a neuron. And inspired from a neuron, an artificial neuron or a perceptron was developed. So if we focus on the structure of a biological neuron, it has got dendrites and these are used to receive inputs and these inputs are summed up inside the cell body and using the axon it is passed on to the next biological neuron. So similarly a perceptron receives multiple inputs, applies various transformations and functions and provides an output. As we know that our brain consists of multiple connected neurons called neural network. We can also have a network of artificial neurons called perceptrons to form a deep neural network. Let's understand how a deep neural network looks like. So any deep neural network will consist of three types of layers. The input layer, the hidden layer and the output layer. So if you see in the diagram, the first layer is the input layer which receives all the inputs. The last layer is the output layer which gives the desired output and all the layers in between these layers are called hidden layers and there can be n number of hidden layers thanks to the high-end resources available these days and the number of hidden layers and the number of perceptrons in each layer will be entirely dependent on the use case that you're trying to solve. And there is mechanics to decide the number of hidden layers. However, we'll not get into that in this session. Now since you have a picture of deep neural network, let's try to get a highle view of how deep neural network solves a problem. For example, we want to perform image recognition using deep networks. So we'll have to pass this highdimensional data to the input layer. And to match the dimensionality of the input data, the input layer will contain multiple sub layers of perceptron so that it can consume the entire input. And the output received from the input layer will contain patterns and will only be able to identify the edges and images based on the contrast levels. And this output will be fed to hidden layer 1 where it will be able to identify various face features like eyes, nose, ears, etc. Now this will be fed to hidden layer 2 where it will be able to form the entire faces and sent to the output layer to be classified and given a name. Now think if any of these layers is missing or the neural network is not deep enough then what will happen? Simple we'll not be able to accurately identify the images and this is the very reason why these use cases did not have a solution all these years prior to deep learning. So just to take this further we'll try to apply deep network on an MNEST data set. So the emnest data set consists of 60,000 training samples and 10,000 testing samples of handwritten digit images. And the task here is to train a model which can accurately identify the digit present on the image. And to solve this use case, a deep network will be created with multiple hidden layers to process all the 60,000 images pixel by pixel and finally will receive an output. So the output will be an array of index 0 to 9 where each index corresponds to the respective digit. So index 0 contains the probability of 0 being the digit present on the input image. Similarly, index 2 which has a value of 0.1 actually represents the probability of two being the digit present on the input image. So if we see that the highest probability in this area is 0.8 eight which is present at seven index of the array. Hence the number present on the image will be seven. So this is how the handwritten image processing happens. Let me practically execute this use case for you. So this is my PyCharm IDE. First of all, let me show you the data set. So this is my MNEST data set and it has got four GZ files which gets extracted when my program gets executed. Now the program or the deep neural network using which I was able to create a model to process all these images and train my machine is this create model_2 py and it's a good lengthy program. So I'll not be explaining the entire program for you but let me tell you the technology or the framework with which I was able to implement this. So I've been using TensorFlow which is one of the open-source Google libraries for deep learning. And right here I have imported tensorflow and then I'm using this emnest data and finally going ahead and creating a deep neural network. So these all things are here. It is creating a deep neural network and the hidden layer that is required to process all these images. And finally I'm creating a model and I'm saving this model with this name right here model 2. CKpt. Now if I run this code it is going to take a very long time. So give it some time. It has extracted all the files and it has started its training. It is at step zero now. So in order to completely train this model, it is going to take 20,000 steps. Let me show you in the program as well. So here it is. So in here I've set the steps to 20,000, but you can always configure it to a number that is,000 2,000. However, you'll have to run this code for n number of times. so that you can achieve a particular accuracy. So after executing for 20,000 times, what happens is a model is created with an accuracy of 92%. So what does it mean? It means that if you pass a particular image, out of 100 images of the model, 92 predictions will be correct. 92 of times this model will be able to tell you the exact number that is present on the image. Let's now wait for this program to execute completely otherwise we have to wait for ours. So I've already executed it once and the model is already created. So what is a model? A model is nothing but a set of files and these three files along with the checkpoint files. Now there is another code which is predict_2. py in which I'm restoring the model and let me show you the line where I'm restoring it. So it's right here saver.restore model 2.CAPT. So this is the name of my model. So I'm restoring my entire model that was created after training of 20,000 steps on MNEST data set and I'm passing an image that is 7_o.png. So it is in this folder in the test folders. I've got other images and I've got here 7_o. Now this is an image. It's the name of the image. I'm not telling the program what is the number. So I'll just stop this training now. And now I'll execute the prediction part where I'm restoring the model. And this model will tell me the number the handwritten number that is there in the image. So this was the image 7_o and the prediction for this image is seven. So my model was accurately able to identify or predict the number that was there on the handwritten image. So let me change the image now and let me execute this again just to show you the image. All right, there's one good question that I would like to take. So Akil asked that how are you saving the weight for the neural net? Can you show us? shorter. So in the previous file that I executed, I showed you that I'm using an object called saver and using the saver object I can save the entire model and weights are also automatically saved along with this model in the checkpoint file. So using checkpoint we can actually reach the final state of the training and then we can use the prediction model. I hope that is fine. All right. So if you see this seven, this is a handwritten image. This is somebody who writes seven like this with a strike in between. And now I'm passing a different image of seven. It's 7_1. So it's different from the first one. And I'll run this code again. Now this time the seven is different and my machine learning model should be able to predict that this is a seven because people write seven in different ways. Somebody likes seven like this. Somebody writes seven and makes it look like a one, they make the top part very small. So there are different ways of writing seven. So however, a machine learning model should be capable enough to find out that as well. So let me just close this and see the prediction. Our model was able to predict this seven as well and predict the value seven as well. So let me execute it again for you. So both the sevens were different but still the prediction is correct for both of them. So now let us go back to our presentation. So after the mnest application let me show you a few more applications of deep learning. The very first is face recognition. Let me give you an example. So all of you are using Facebook and you do spend some time on it. So if you remember a few years back when you used to upload pictures with your friends, it makes a box around a human face with a box appearing at the bottom to ask you to type the name of the person to tag him. So it was able to identify that it was a human face. But now it is able to auto tag. It is not only able to detect faces but also identify who it is. And how is this possible? It is only possible using deep learning. And Facebook also has a deep learning library called cafe 2 using which they have applied all these things. The next use case that is implemented using deep learning is Google lens. This is one of those applications that has been recently launched for smartphones by Google. What does this app do? You just have to install it, open it, point your camera on a particular thing like this flower over here and in real time image processing happens and Google will get back with the entire details of the object like the name of the flower where it is found etc etc. So if you point it at a building or any shop it will tell you what kind of a store it is. If it's a restaurant it will show you reviews the ratings menu etc. So what is happening here is that in real time you are able to use a deep learning net and get all the information you want and these applications are really amazing because it directly brings deep learning to the end users or the common people. So they can easily use the benefits of deep learning without worrying about what is happening at the background. The next use case is the machine translation. This is again a very important use case. And there is also an app in play store and this is called translation app. So here is an image that says more chocolate. I don't know what it means. I don't even know which language it is. But with this app what you can do is that you can capture the picture of the packet and this app will first detect the text in the image then extract the text like this and then translate it for you. So for example it has detected the text extracted it like this and here it has translated morg which means dark and then it writes back again on the image. So what is happening in this particular use case is that first an image is captured. Image processing takes place. Text is extracted through processing and once we have the text we translate it to the desired language which is English in this case and then again image processing happens where we are writing the text on an image again. So more chocolate means dark chocolate. So this is a really great use case because this is a combination of multiple learning algorithms like CNN and RNN. So these were a few more applications of deep learning. I hope that you found them interesting. So first thing which comes to our mind there have been lot of emphasis on this term called artificial intelligence. So let's first try to understand that what is artificial intelligence on a very high level and why we may need it in first place for solving a problem. So let's try to understand with an example. Person goes to a doctor and he wants to get checked at whether he got diabetes or not. And what doctor would say is okay there are some tests which you need to get done and based on the test results doctor would have a look and from his experience from his studies and previous examples patients he had seen he would be able to evaluate the reports and say that the patient has diabetes or not. So if you just take a step back and think I said that doctor has experience. So what do we mean by experience? The doctor has learned what are the characteristics of somebody having diabetes. Will it be possible if we can provide this experience in the form of data to a machine and let machine take this decision whether a person has diabetes or not? So the experience which doctor learned through his studies and his practice what we are doing is we are taking customers data who have with different reports and different parameters on different things like the glucose count in blood or the weight and height and all these parameters about a human being and based on that we have fed it to a machine and tell that what are the characteristics of a person who has diabetes and from this let's say we have 1 million customers data we have given to a machine and let machine do this stuff from his experience which comes from the data or historical data to be precise and do the same task which a doctor is doing. So what we have done is if you see from this example what an artificial machine or artificial intelligent machine is doing that it's learning from the historical data and trying to do the same thing which an experienced and intelligent doctor was doing. So this kind of area or domain activities which human beings were doing if we can make machine intelligent enough to do the same task. Why we should create these artificial intelligence based machine systems on a very high level there may be lot of points but if you just discuss a few points that human beings have limited computational power and we guys may be good in terms of classifying things like you know you can see your friend in a group photograph and easily can say who is your friend and who are others. You can easily listen to a language and comprehend what a person is doing. But human beings are not very good in doing lot of mathematical. If you try doing good amount of mathematical computations in your head probably it not be very easy. And second is that it's not possible for human beings to work continuously let's say 24 by 7 a day for 30 days continuously. If we can make machine do such stuff, one they would be able to kind of do these computations very fast and like we spent a good amount of time in discussing the GPUs and I also mentioned that Google is talking about a TPU machine transfer processing units which would be hugely changing the entire paradigm of computations and machine would become more and more competitive or even better than human beings in some of the fields. This is the formal definition but if I loosely translate it's basically that artificial intelligence machines are those machines which can do task which human beings can easily do. So things like identifying what's written let's say in a license plate or playing games and I'm sure some of you would already heard that machine have defeated the go champion and the chess players. uh now we have digital agents like Siri and others which can understand what we want them to do and can take intelligent decisions from the text or from the voice itself. Basically these are very high level and some of the fields where deep learning is made a great great inroads something like game playing expert systems self-driving cars robotics natural language processing so there may be different and new areas where we are implementing all of you know that everything every experience of human beings is getting digitalized the kind of things you buy kind of things you watch and what your preferences are who do you like on Facebook who you don't like what kind movies you like and all these things in terms of reviews being captured online. So once your data is going and captured online there are systems which can analyze this data. So given this huge data generation as well as now we have machines which can process it and make some intelligent decisions are available. So that's why you will see there have been lot of emphasis now in last couple of years lot of new things are coming in. Some of you who have been reading these papers on different subjects, different architectures would know that most of these architectures are not very old. It's a very dynamic field every day. In fact, on a weekly basis, you will be hearing about a new API or a new kind of architecture being developed by somebody. Most of the stuff we will be studying in these classes are not very old like convolution neural network and recurren neural network. Some of their variants are as new as as last year. If you guys follow TensorFlow closely, they introduced a library called object identification. Object identification, object detection API which TensorFlow has made available for everybody. You would be able to see it to yourself that this API works. There are five different options of selecting different deep learning architectures or convolution neural networks for this API. But it's been able to identify human beings and all other 90 objects there with almost 99% accuracy. In some cases from even human beings would be finding it difficult to kind of see and predict what the object is. But this machine has gone even beyond a human capability in terms of identifying objects. Given that we have a fair understanding or very high level understanding we haven't got into details of artificial intelligence but basically from a loose understanding that artificial intelligence of making decisions or machine making decisions which human beings were earlier doing the task something like game playing and natural language and driving of car. Let's understand how this machine learning and deep learning are related to artificial intelligence. Given this learning that now your machine is able to understand and learn from the data, we can solve multiple business problems with the help of this. So let's take a very small example. It was like whether a person has diabetes or not. And I was mentioning that this kind of decision being taken by the doctor based on the reports he has got. And these reports have some numbers like number of time a patient had a particular kind of issues or what is the glucose count and what is BMI and what the person's A is and BS of these numbers the doctor was able to make this kind of decision. We can take the same analogy where we were trying to predict which species of FL it is. We can take information of patients and different attributes on different features of a report and the patient and the machine would be able to learn from these data sets and for a new customer or a new patient it would be able to classify whether the patient has diabetes or not. So there are two sections of it. First one is the information about the patient and different characteristics of his health. So from this which is number of time and glucose count till age these are the information points about the patients and the last column is the information whether the patient has diabetes or not. This kind of problem where we have some information points where they explain what the situation is and other in the last column or the information of output is in some kind of classes. There's a specific type of machine learning problem it is but as of now the characteristic is that we have some information about the patient and the last column is telling me whether the patient has diabetes or not. So what a machine basically learns it that it learns all those rules in the example which I was quoting that earlier cases people used to create these handcoded rules to predict whether an event will happen or not. But in machine learning, your algorithm will learn from the historical data and see what are the combinations which decide whether a patient has diabetes or not. And these combinations would be of something of this type. It is only for illustration. It's not the real numbers. But it's for illustration that your machine or your machine learning algorithm is being able to identify these rules based on the historical data. So after learning it has created the glucose count is less than 99.5. If yes then go to next one. If no the person does not have diabetes. If glucose count is greater than 166.5 yes the person has diabetes and if no then there are further drill down of rules. So all these combinations or rules are dynamic in nature and what I mean is that these rules would be changing if your data says changes and you can take the same model can do the work whether a patient has diabetes or not. And you can take the same model and make it learn on a new data set. Let's say flower species. It would be able to learn the new rule from itself. So the intuition like human beings were learning from examples. Your machine learning algorithms also learn from examples. But just to frame our problem statement that machine learning we know it learn by experiences and from the data from the historical events. There are three kind of problems which may be interested in solving. First one is called a supervised learning problem and supervised learning problem is basically occurs when you have some input variables and one output column. So both the examples which we discussed till now one was on the flower species where we are taking data on different features of a flower and then which species of flower it was. So the last column is the dependent variable or output variable we are trying to predict and all the information variables are called input features or input variables. So input features or input variables it's kind of interchangeably been used in different texts. Input and output these are the two different sections of supervised learning. Why it is called supervised learning? Another take if I need to explain it that we have a column to guide the algorithm whether it's making the correct decisions or not. So let's say your model says person has diabetes but the actual data says the person does not have diabetes. So you have some kind of correction mechanism within your data itself which can help your model tune itself to make better predictions. So this kind of output variable some text it's also been called a teacher variable. So it's guiding the algorithm to decide those rules which I just go through. Another type of machine learning is called unsupervised learning. And in unsupervised learning we have only the input features and our objective is that we should be able to identify the patterns within the data itself. So some of you who are working in telecom domain or marketing campaigns you would be very much familiar with segmentation analysis or cluster analysis where our objective is to identify coherent groups within the larger population. We take the customers as it is the whole population and based on different parameters and variables we identify some of the groups of customers or products whichever the business problem is to identify which are similar in nature so that we can take either marketing campaign or develop new products for those specific groups. Final is called a reinforcement learning. A reinforcement learning is a kind of learning where the agent learn from the environment. So it works kind of reward and penalty. You can think of a self-flying helicopter. So you leave it in the environment and it'll be deciding based on the wind speed and other parameters in the environment that how much it should fly and the reward is that the fuel should be efficiently be spent and more time it should be spending in the environment. So supervised learning as I said that the objective here is that we have some input features and input features would be holding information about different aspects of a given problem or customers and we have an output variable which would be explaining whether the event happened or not some kind of output variable. So there are two types within supervised learning. One is called as regression. Second is classification. And the differentiation happens only because of the type of output column. If your output column is the type of numerical values or continuous values like numbers. So that would be a regression kind of problem. A very high intuition level. For example, you're working on a problem where you need to predict how much would be the sale of your company given the information that how much they're spending on marketing, how many employees are working, which month it is of the year. And if you have this information, you are going to predict what is the million dollar of sales your company would be doing. So these kind of problems where your output variable is of numerical values, then it's a regression problem. On the other hand, if the dependent variable is of categorical nature or of discrete values, it signifies that it's a classification problem. Given that it's a categorical values, your objective is that how you can put the different customers or products into different classes. So that's why the name suggests classification problem. So as I said there are two kind of supervised learning problems. One is regression and another one is classification. So let's take a use case where we need to predict the housing price of a particular locality. And we have information about these houses on different parameters and these parameters are like these. So let's say what is the crime rate in that area? How old is the home? The distance is how far it is from the city. This is from Boston housing data. So this is if I'm not wrong it's percentage of black population or some variable. We have the description later on and what is the actual price of the home and all these features from crime to isat is the information about the house. So these are my input features and the output feature is the price of the house. This is in million dollars. And our objective is that we should be able to fit an algorithm that it should be able to learn from all the historical homes which were sold based on all the features and what was the price it was sold for. And once the model is trained, it should be able to predict that how much should be the price for a given home. Let's take an example. Let's say one of you is interested in buying a home in the Boston area and you would like to know that what is the ballpark figure for a two-bedroom flat which is of some square ft and let's say 20 miles from a specific location what should be the price. So one way would be you go and talk to people and try to understand that what has been the average price or if you have this kind of algorithm available which can help you understand that given these features that was the price and if you can create a regression model it would be able to help you that given some features of the new home which you are interested in buying what should be the price of it. So as I said there are two sections. One is the independent variable. So all these information about the home and the last variable is dependent variable which would be information about what was the price. You see this is a kind of scatter plot between the distance from the city and price for the home keeping all other variables constant. We are not looking at the influence of other features. But we are looking if we just need to model or if we need to find a relationship between the distance from the city and what is the price from this graph you can make out that further the city houses the lower the price would be if you keep all other things constant. So here it's like if we can identify this kind of relationship that's called linear regression. But for a given distance from the city you would be able to predict what should be the ballpark figure for a house. If you just have this information not all other information which we have talked about in a similar fashion how it's been done is that there would be a relationship between the price of the house and all the features which we discussed. So this is only a relationship between one of the variables distance to the city and the price of the home. In a similar fashion we would be able to find relationship between the price and all other features. So all of the features if you know like how old is the home, how big it is, what is the crime rate and all. So this kind of model is called a regression model and it's a very basic equation of a straight line. Y here is called the dependent variable. A is intercept and B is called slope and X is called independent variable. If you go deeper and try to understand what it's basically doing is this equation is trying to tell me that if I already know the relationship if I know the value of a and b from my historical data which is about different homes given the value of x I should be able to calculate the y in our particular example is price of the home. Let me try given an example what slope means. So slope is the change in dependent variable. If we change x which is the independent variable by one unit. So let's say if I change x by one unit, how much change in happen in y? Help me understand what is the kind of relationship between x and y. And a is the value which tells the value of y when the x is zero. And you can think of it something like that. If you put the value of x equal to zero, whatever the value is y, then that's the intercept. But basically from intuition perspective, you can think that an intercept is the value which is there even though you don't have any information about x. For example, we were discussing the relationship between the distance from the city and the housing price. Even though the house is exactly in the city, then there would be some value and even though the house is 100 miles from the city, there would still be some price. So it help us kind of intuitionally understand the relationship. It also help the line to understand where it start whether it start from the origin or some place within your axis. And this kind of equation is called equation of linear regression because if you see here the power of x is one. So that's why it's linear in nature. And it is also that we are fitting the relationship between x and only one of the variables. Multiple regression where what we do is that instead of finding the relationship between only one variable and the dependent variable in most of the practical scenarios the dependent variable Y is dependent on more than one features. So for an example, your house price is dependent on all these features, all these information points available and all you want is that your regression model would be able to identify a relationship giving all the information together and then predict what is the price. The equation becomes y= to b1 x1 b2 x2 or b3 x3. this kind of equation where B, B1, B2, B3 and all these coefficients help us understand that what is the contribution of a single or of a given variable into your regression equation. So let's have a look that how you can fit this model in Python. So if you look at the first block of the code where we are saying import panda as pen pd, import numpy as np and import num metplot lib as plt. This is the convention in Python to import some of the libraries which we'd be using. So these are the libraries which are required for running this module or this this regression model. So once we import these all the functions available in these libraries we can call them very easily and we will be seeing that how you can call them. So once you have imported these libraries if you see that we are loading the data called Boston and that this is the same data set which we have been discussing in terms of a use case here. So next line of code so here we are importing the data and loading the Boston data set. The next line of code we are calling the pandas library because we have imported the pandas library as pd. Then we are calling a function called dataf frame so that we can you know create a data frame in python and we are creating it Boston data. So it will be creating boss as a data frame. So data frame on a very loose term you can think of kind of spreadsheet kind of format where your data is being put in rows and columns and you can think of an excel file kind of framework for a data frame though it will be different but just for intuition purpose and after importing I'm calling dot head. So what do head does it will be giving top 10 rows of my data set. So there are all the 13 columns in python index start from zero. You can see that index started from 0 1 2 3. So you can see what this data is. This line of command which says dot columns. So dot column gives you all the features available in your data set. So these are the different feature names. So these are the different column names for the data and the price. In the end of the code, I have written actually one line of code which can give you all the details that what the target variable is, what was the history of data, where it was recorded and all. So you can easily look at this. We are calling this Boston.target and we are calling it as a boss.pric. So we creating a column in our data frame which was boss from Boston.target. So there is another data vector available in the Boston data set itself. And now specifically we are saying y is equal to this particular variable y we will be representing our dependent variable all the features plus the Boston price. So boss dot drop price x is equal to 1. It means that we are dropped the price variable from the overall data frame and xis one specify that we are removing the column. So x's 0 represent the rowle operations in python and x's one represent the column level operations. Print statement we are just printing now the x. So this x is all the input features of our data set. So all these columns which we will be using for predicting the housing price and how it will be working. So it's not an actual model but what actually happening is once we have created the model it will be doing that it'll be fitting a line which would be going through the actual data set would be something like this that price is equal to some intercept term plus b1 multiplied by crime b2 multiplied by another variable zed n then b3 multiplied by another variable and so on and so forth and this intercept and b1 b2 and bn would be the coefficient which your model would be learning from your already available data and here we are showcasing top five values of our housing price. So y is the dependent variable and we are looking at what is the top five values. So this was only a very brief and very basic introduction that how do you import our data how you can see what are the different columns. It has nothing to do with machine learning but it is only for people who are new to Python and for people who have been out of touch in Python if just want to brush up skills and this line of code if you have a look which I'm highlighting now it is we are using a scikitlearn model for test train split and what it's doing is for both because we have already x and y the test sizes we are saying 33 so it's basically we are randomly selecting 33% of the data for we putting it sep separate in the test bucket so that we can test it later on. And this random state five means because we'll be randomly selecting if you specify a random state. Every time you run this code, you will be selecting the same set of elements from your data set. It help you understand that the variation if you run the code multiple times by changing the variation is not coming because of the selection of sample. It should be because of the different model changes you are making. This dotshape function in Python specify that what is the dimensions and if I mention the first one X train.tshape is giving me 339 and 13. Basically it's telling me that there are 339 rows and 13 columns in the data set. X test there are 167 rows and 13 columns. And your Y test is just 339 rows and there's just one column or it's just one vector. The number of rows in X train and Y train are same. In X test and Y test the number of rows are same because they have been selected for the same combination. So same houses we have selected the input features as well as the corresponding values of output and same has been done for the test section. What we are doing here as I was saying that we have imported a library called scikitlearn and scikitlearn has different modules for different machine learning algorithms and linear regression is one of the modules in scikitlearn. So we can call this scikitlearn module from linear regression called lm and this equal to and in python it's called assignment variable. So we are assigning lm as a linear regression module in the scikitlearn. And now what we are doing is if you look at this line only lm.fit. So basically we are telling that use the linear regression module from scikitlearn and fit the model between x train and y train. So basically what we are telling the model that you learn those coefficients for different x values given y values in the training data set. Basically what the fit function does it calculate the values of your intercept B1 B2 B3 for all the features in your input features for a given Y variable. So once we have fit in the model and once that has been fit we can use the same model for doing the predictions. So let me remove it. It should be like this. So LM.fit fit we have fit in the model and once there has been fit we can use the learned model which is lm with a function called predict x train. So what it will be doing is once it has learned those coefficients B intercept and B1 B2 B3 for all the features and input you can use the predict function for making the predictions for your training data set and you already have actual values as Y train and then you can compare that how good your model is doing and how you can compare it that if you look that we are put together same thing we have used for the X test data set LM.predict predict X test and again the prediction has been done. So if you see here I have put together as a data frame Y test and Y test bread and the difference look like this for the first value which was 37.6 and the actual value was 37 this value then the predicted value this actual value is this and this difference between the actual value and the predicted value signifies that how much is the error in your data set. So had it been that your model is giving the same prediction as it was the actual value you would say there is no error your model is 100% accurate and all the predictions being made by the model is absolutely you know bang on but normally doesn't happen you end up having predictions which are a bit off from the actual value and we measure the difference as one of the characteristics or one of the parameters to identify how correct your model is. In this particular statistic there are two metrics being used for identifying but the most uh basic one used is called mean squared errors. And what mean squared error is it is basically the difference between actual value and the predicted value by the model. And what I mean by this is that let's say this is your predicted value 37.6 and this is your actual value. What you do is you take the difference of these two and then take the square of it. Why we take the square of it? Because this in some values the difference may be negative or positive and if you sum it up the difference may come to zero and you may end up thinking that okay model is doing really good stuff. To avoid it, what we do is that we take the square of it so that the difference between actual and the predicted becomes positive and you can sum it up to showcase that how far your predicted values are from the actual value and then you take a mean of it to showcase that what is the mean difference between the actual value and the predicted value. It can also be used for model comparisons. Here I can show you that how it is working that let's say you have some actual values something like this and let's say you fit a model I fit a model so there is one model prediction one another model is prediction two so what you can do is you can take the difference between the actual and the predict so 10 minus 2 is 2 and then you take the square of it which is four 23 and 21 again two square of 2 4 then third one the difference is five and then square of it is 25 and so and so forth for all the values You get the total value of sum of squared errors and you divide it by the number of inputs which is five and you get the mean of the squared errors and you do the same thing for the second one and if you see it is very less five. So probably it would be able to help you understand that which model is doing a better job in terms of predicting the housing prices or any other numerical variable. And there is another statistic which has been used for identifying how good your model is which is called mean absolute percentage error and that's basically the absolute difference between the actual value and the predicted value absolute terms and you sum it up all the values for all the entries and divide it by number of all the value of absolute value of your actual values and it can help you understand that what is the average percentage your predicted values are different from the actual value. So sometime if you see it will be somewhere in percentages. So what I've done is I've taken the absolute difference between this value and the predictions. I sum it up and divided by sum of my input values. So whatever value comes in you can say okay it's 5%. So it'll be fair to say that your model is 5% off from the actual values or the error term in your model is let's say 5% or 6%. And whatever predictions you're making from the model, you can keep a buffer of that percentage when you share it with the team. And what I mean is that let's say your mean absolute percentage error is around 10%, and it's about the sales of a company. So when you share this forecast, you say that my predictions are around 90% accurate. They may be actual sales maybe plus - 10%age. So this can help you giving this kind of variability in your predictions. However you have implemented code in Python itself or the scikitlearn library, you can call mean squared error the function from skarn and it can help you calculate the MSC for a given model. So basically there was a very quick introduction to linear regression though there are different applications but one thing remain common that we are trying to predict the dependent variable whose nature or the type of dependent variable is a numerical or continuous data. Some of the applications like predicting life expectancy based on these features like eating patterns, medications, disease etc. You can predict housing price. We have already seen the example on that. We can predict the weight on different features like sex, weight, prior information about parents and all. And you can also predict the crop yield of crop based on different parameters like rainfall and all. And as I said this like very limited uh use case uh list. I'm sure people who are working with sales department you have to make predictions how much would be the sales. People who are working with call centers you need to predict what would be the number of calls for next month. People who are working with the marketing you need to predict what would be the footfall in a given company or a mall. So there are different applications of regression models but one thing is common across all these applications is that the dependent variable which we are trying to forecast is of continuous data type. So let's get moving to the next agenda for logistic regression. So at the time of the introduction to machine learning, we discussed there are two kind of supervised learning techniques. One is regression and other one is classification. And the major differentiating factor between the two were that in regression we held a dependent variable of continuous values and in classification problems we had a dependent variable of categorical types. So let's take an example of how we can do it. So here let's take a use case where we have got some information about some customers and the data set looks like this that we have some customer ID or user ids gender of the customer or the user his or her age estimated salary every month. So you can think of in any one of the currency either INR or dollars and whether this user purchased an SUV or not. So as I was saying earlier the dependent variable here is 0 and one. So it's a discrete value or categorical value which we need to predict and the features which we'll be using in the model are age and estimation. Why we would need a logistic regression kind of algorithm? It would be a straightforward process that if I take purchase as a numerical value 0 and one and I take some input features like age and estimated salary and you will be right in saying to some extent that this is a possibility of doing it. So there are two major problems coming if we follow this and some of you can help me what may be the problems if I try using the linear regression for solving this kind of problem. But one limitation I can think of is that here I'm looking for an output which can give me some kind of probability that how likely I am to buy a product or service. So one thing the limitation or the restriction with probability is that the probability term should be between zero and one and zero signifies that there is no probability or there is no likelihood of event happening and one means that it's certain that the event will happen. There is no possibility that we can have probability values less than zero or greater than one. So if I'm fitting a linear model taking the purchased column as my dependent variable, my values because the linear regression has no such limitations can pass these values beyond one or less than zero. So what I require is that I fit the model in the similar fashion like I did the linear regression the equation I used earlier. that I want that information would be coming from my features in the similar fashion. But what I want actually is that this y should be mapped to the values between 0 and one. And given the limitation we have just talked about probability that it should be between 0 and one. It should not go beyond one and less than zero. I need to find ways if I can kind of force fit or kind of force this y value which would be coming from this equation and I force fit into a values between 0 and 1. So to solve this problem there was a function called sigmoid activation function which would be extensively be used in our uh deep learning as well at different places. But logistic regression comes from this activation function itself which is a function looks something like this that output value would be 1 divided by 1 + e ^ minus x and x is not actually the one of the input but any value we are giving it. And if you fit any value into this particular equation it can convert any value between minus infinity to plus infinity. it will map it to between 0 and 1. If your value of x which you are putting in here I could have selected a different value different name at least but if you give the highly negative value the output would be very very close to zero. If this input is positive then the output would be close to 1. If the value of x the input here becomes zero what would be the output? 1x2 because any values power 0 is equal to 1 and 1 / 1 + 1 would be equal to 1x 2 or half. So logistic regression is nothing but an extension of your linear regression itself with one additional fact that you want to force fit your output between zero and one and for that you are using activation function called sigmoid activation function or sometime it's also been called logistic activation function to do the same task. So this is an intuition behind your logistic regression where you take values of your equation from intercept and different coefficients for your input and you map these outputs between 0 and one. So once we have understood that logistic regression is nothing but the extension of your linear regression only with a restriction on the output being mapped between 0 and one. We are shifting had we are fitting the regression equation we would be having scenarios where the value would be going beyond one or less than zero and to avoid this scenario we fit in logistic regression with the help of sigmoid activation function which looks like this and if you see as I was saying when your value of your model go beyond let's say this is r0 so all the values which are positive and greater than zero the curve goes in tangential towards one and for all the values which are less than zero it goes towards zero and at the place of zero the probability is 0.5. So it's a 50% probability if your output is very much close to zero or it's zero. It can be used for multiple scenarios. One of the example we are taking is the example whether somebody will buy an SUV or not. But if you're trying to solve problems like somebody will say yes or no to a product or service or whether something is true or false or high low or any different categories but logistic regression can easily be put in for multiclass classification problems and basically if I just give you a very quick introduction how it works is that in multiclass classifications it kind of does mapping that one class versus rest of the other classes and then same analogy follows that which class a particular ular event would be associated with but end of the day for whichever class or category the probability is highest the model will predict that uh it should be belong to that particular class like MSC we have a statistic or a parameter to evaluate how good your model is doing and that was a parameter to check that what is the difference between the actual value and the predicted value and how we were doing it we were taking the value which was actual subtracting the predicted value taking square of it. Do it for all the examples and divide it by number of training examples we have and then it gives you some number and I was also saying that this number is helpful in kind of comparing different models. So let's say you fit a model I fit a model and we compare MSSE for both of them. Whichever model is giving me a lesser MSSE it is kind of an indication that probably your model is doing a better job in terms of prediction than mine. In a similar fashion, we needed a kind of a statistic to see how good your model is doing when your model is doing a classification problem. So here there are four categories that let's say we have only two classes good and bad actual values good and bad and what your model is predicting good and bad. So four examples which belong to good category and your model is also predicting them good category. So this type of events or examples are called true positives because your model is doing correct prediction on positive examples. Another category which is your actual value for those examples is bad. They belong to bad category and your model is also predicting them bad. These are called true negatives and these are correct predictions because whatever the actual value is your model is also predicting the same thing. However, there are two categories where your actual value was bad but your model is predicting good. These kind of examples are called false positive because your model is falsely predicting then these are good examples. And another category or last category is called false negative where actual value were good and your model is predicting bad. So how do you learn or how does your model say that which model is doing good job. So what we do is we calculate what is the percentage of values examples have been predicted correctly. These sections in blue true positive and true negatives these are the examples which your model has been able to predict correctly. And these two groups false positive and false negative are the incorrect predictions. So what we do is we just want to take what is the percentage of correct predictions. And this matrix is also called confusion matrix. Let's say there are some examples out of which 65 examples were there where actually they were good category examples and your model is also predicting them as good class good category examples. 44 are those where they belong to the bad category and your model is also predicting them bad. This is 44 and eight are uh actual bad and prediction is good and four are actually good and uh predicted bad. What you do is you sum up all the correct examples 64 and 44 and divided it by all the examples in your data set all correct and incorrect ones and here you get 89%. So all you can say that your model is being able to predict 89% accurately or if you want to explain it to your business team and say that whenever I give you a prediction that uh 100 customers will churn and I give you a list of 100 customers I can say with certaintity that at least 89 will churn from them with some certaintity. So because your model has given you 89% accuracy. So that that's how it's been kind of communicated to business teams that we are thinking that our model is 99% accurate and whatever prediction we are giving we are very very certain. But if your model accuracy is 70 or 60%. Then when you give the predictions to your business team you say that okay though we are giving you the predictions but we are not very certain whether it'll work correctly or not. So this accuracy percentage is in a similar fashion like we did for linear regression as MSC to identify how good the predictions are. Your accuracy percentage is another metric to see how close or how correct the predictions has been. So now we can see the implementation of logistic regression in Python. So first few lines if you see we are importing the libraries or the machine learning libraries which we require to do the data manipulations. We are importing the data which is a CSV format and this data is already available on your LMS. If you want to import you can easily import from the LMS itself unlike the Boston data set which we were importing from the library itself. Here we have got a flat file as social network adscv and you can call read csv function of pandas library to import the data. So you are importing the data as data set and as I said head showcase the top five rows of your data. So here we have only five columns. One is user ID, gender, age and salary. And the last column is our dependent variable which signifies whether a customer or a user bought the SUV or not. So it's 0 and one and one means the person bought. In the previous code, we used one convention of selecting X and Y. Here we are showcasing another way of selecting that's called eyelock. So we are looking for the location and this convention if I go through what we are doing here that this is the data set within the data set we are specifying the locations this colon means that we want all the rows and as I was saying earlier that in Python the index start from zero. So what we are saying we want column two and three. So what we mean that this is zero this is one this is two and three. So we want as our input features two and three and the values. So it'll be creating an array of these two columns. We could have used gender but I will leave it to you that first we need to create the gender as a vector of 0 and one. So you can create a function which will say okay if gender is equal to male then one else zero or you can create dummy variables there are function available in scikitlearn. So it's an exercise for you that this is the code already available but I would encourage that if you can also include gender information into your model. The next line of code why we are saying the dependent variable is all the rows and column number four. So column number four is your purchase information whether a customer bought the SUV or not and again the values to convert into a kind of list format. So we have specified two things. The two columns is the information about the input features or the information about the user in terms of how much money they make and what their age is and uh information of why whether a customer bought the SUV or not. And the next line we are doing the train and test split for the same stuff to evaluate whether the predictions been made by the model on the training data set on which the model was learned is still doing the correct classification on the data set or the test data set which was not involved at the time of training and this 0.25 means that we are selecting 25% of the data for test and remaining 75% for a train. What is the correct split of train and test? Normally it is correct to choose between something like 6040 or 7030 or 80/20. If your data set is big enough then I think having 80/20 kind of split is good or whatever you can try these different combinations but as a rule of thumb most of the time I have seen people taking something like 6040 or 7030 kind of distribution between actual value and the predicted value. Now there is one important thing for data prep-processing and this selection which I have made for doing the data prep-processing and some of you who come from the machine learning background will already know that how important it is to kind of scale your data and what do I mean by scaling that if you look at the data set which we are using for input one is the age column and second one is the income column age can be somewhere between let's say 1 to 100 or 120 at max and your income is in like some thousand and some 100,000 numbers. Both these values are on different scale. Scaling your data on let's say all the values between 0 and 1 will help me understand that what is the importance of each variable. For for example, if you look at a regression equation and you see those coefficient B1 B2 for all the input features these feature or these coefficients can give you kind of indication that how important a particular variable is. But this intuition will only be correct if all my features were on the same scale. If these features like age and salary when they are on different scales, you will not be able to compare what these coefficient really mean because there are two different scale your values come from. So it is always a good idea to have all your features on the same scale. There are multiple ways of doing it and there are multiple type of scaling parameters. The simplest one is called minmax standardization. And what does it mean is that for a given column let's say we are talking about age column which is 19 35 26. If I need to do minmax standardization what do I mean is that I take the value it is let's say 19 and minimum value here is let's say 19. I have only these five values. So how it works is that this is the formula. This is the value or how it's being presented. X I minus the minimum value of the column. So let's say age divided by max of age minus minimum of range as well. But basically what this formula will do if I do it for all the values in the age column, it will be converting all the values between 0 and one. And there are other ways. I also said that there are normalization process which is like you take the value minus the average value divide by the standard deviation if I call it correctly. So whichever method we apply all I'm saying is that these values of age and salary should be brought to the same scale. So if I'm applying this minmax standardization I'll apply to both my columns so that both these variables are on the same scale and I should be able to use them in my model. And this is again a very important thing that whenever you do a standardization you will be using this process that you fit the normalization or standard scaler on the train data set and you use the same learned standardization from the train on the test data set. But basically how it helps that your data set on which your model is being trained it will be converting the values between 0 and one based on the minimum and maximum values. If the test data set have different minimum and maximum values, it can have different value for the same number. So that's why the process is that we make our standardization fit on the training data set and use the same minimum and maximum value for test normalization as well. It gives the same scale for all the values and for model predictions. It's very helpful in a similar fashion like we called a linear regression object from scikitlearn in the previous example. In exact same way we can call a logistic regression function from the scikitlearn. So it's a scikitlearn linear model and we are importing logistic regression. Now um we are fitting the logistic regression between xrain and y train the same way we did it earlier for the linear regression. And once it has been fit, we can do the prediction for test data set. We are also doing the same thing that we calling the function which was classifier for the logistic regression and we doing it on the X test data set. And here the default probability is 0.5. So what your algorithm is doing in the back end for all the examples wherever the probability in X test became greater than.5 it was tagged as one and for all the examples where probability was less than equal to.5 it was tagged as zero and now we are calling this function called confusion matrix between Y test and Y spread. So we are comparing that what is the values of your true positive true negative false positive and false negative. So this is the values that these are the true positive these are true negative and these are the mclassification values. And if you want to calculate the accuracy you can easily do it by 65 + 24 divided by 65 + 24 + 8 + 3. So all we are doing is we are trying to identify what is the percentage of correctly predicted numbers and uh this is the code of section. So it can do the prediction. If you see what it has done, basically if you look at the section that your regression model has fit this line and you can see it's a straight line and that is why in some of the text logistic regression is also being called a linear classifier. And why it is called linear classifier because it is predominantly being made for fitting a linear equation. The logistic regression equation was y= a + b1 x1 b2 x2 and all these coefficients and respective inputs. But the highest power of your inputs were one. And you would already know that if it's a polomial of power one, it stand for a straight line. So that's why you can see a straight line. There are ways some of you would argue that you know we can fit a nonlinear line with the help of logistic regression. But you would also concede that there are some tricks which we use for creating nonlinear lines through logistic regression. For example, you introduce higher order polomials into your model so that the separation becomes nonlinear. These kind of algorithms are really helpful only solving the problems when the objects are linearly separable. When the separation between the objects is not linearly separable, these kind of algorithms are not very helpful and we need to identify algorithms which can fit in nonlinear hypothesis or separation boundaries between different classes. So let's take a use case to understand that what are the simple scenarios where unsupervised learning can be used and how does it really work. So let's take an example that we have some housing data and housing data in terms that what their locations are and these white dots on the screen in the blue background showcase that where these homes are located and the objective of education officer is that he needs to find a few locations where the schools can be set up and the constraint is that student don't have to travel much. So given this constraint in mind the officer needs to decide the location. There may be easily we can identify if we are not using any algorithm. So let's say if I know that I'm an officer I need to open three schools in the locality and I know the information where the homes are located. I can easily see okay probably this is one location. I'm just highlighting it. And the constraint I also mentioned that student don't have to travel much. What I mean is that if you open the school here then everybody of you would say that it's not a great location for a school given that it's far away from the population. So this is not the correct location and from the perspective of identifying the home probably these three from a human intervention or or like some of you has been given the task without any algorithm you can decide that if you set up the school most of the students would be traveling less to go to a school. So given this problem we can easily see that we don't have a dependent variable as such which is telling us whether it's the correct location or not. All we're doing is that we have number of locations which we need to find schools for and then we have home locations and based on the distance of each home we need to identify which the proper distance proper locations of these schools and another thing which is coming from the same logic that there are no predefined classes of these locations. And one more point if you would like to add and some of you who have done the clustering or the segmentation job in your respective works that these numbers we say three or four or five it's not predefined it is most of the time given by the business that how many clusters or segments they would be looking for though there are statistical ways of identifying that which is the best number of clusters should be but basically most of the time it would be coming from somebody in the business that okay I see that let's say I was working for one of the Indian telecom companies here quite a time back and at that time their subscriber base was around 300 million customers and imagine that if you're trying to create segments for this big a population and if you create three or four clusters you can easily understand that it would be very difficult for marketing team or any product team to design products for such a big population. So though statistically it may look that okay four or five unique segments are there but you end up creating lot of small small segments and there may be a possibility that you will be creating 20 or 30 segments for such a big population. So my intent of saying this number that we are trying to identify three locations within the population has to be decided either by business or people like you who have knowledge about the data as well as that what kind of business they are running and what is the final usage of this segmentation exercise. So let's see one way of doing our selection of these school locations is like we have already doing it. If you identify that somebody looked at the homes and see the densities where the density is high and selected the home automatically but there are algorithms also available to do the task and I can give the name here itself it's K K means algorithm and so first we would like to understand that how does an algorithm work if it needs to identify which is the best location. So if you're looking at my screen, let's say our objective is that we need to identify two locations first and we have some data and it's scatter plot available and we need to identify where the school should be so that the distance from home should be minimum if that's our objective. So how we can do it that let's say we randomly assign two points from the existing data set and actually easiest ways that you randomly pick two numbers from your data set itself and then what you do is you assign these two selected points as these are your cluster centroid. So this is the center of your selected population. And in the second step, so once you have initialized these two random points, then the next step is that you measure the distance of all the homes from the initial selected point. So let's say you do the distance of this home from this selected point and again from this that for each house from these randomly initialized point you measure the distance from the selected point or the initialized point and any home location and see which distance is minimum or which distance is less in comparison to the other from the selected initial point. So we can easily see that this distance is smaller than this distance and this point would be assigned to this particular group. The first initialization step is initialize as many number of centroid as many clusters you need and in the second step you do the cluster assignment and in cluster assignment how it's been done is that you measure the distance from these initialized points and see wherever the distance is minimum and then assign this home to that particular segment or cluster. So this exercise has been done for each home. I'm just trying to show for a couple of them. And based on the distance, the assignment is complete. So this color also signifies what we have done is after measuring the distance for each home from the initial points, we have assigned these points to this cluster and these blue points to second cluster. And then once this assignment is complete, it moves the centrid. So what it'll be doing is it'll be taking the center of all the selected points and then it'll move the centrid from the previous point to the next point based on the new assignment which has already been completed and then what's been done is the same exercise which was done earlier in terms of cluster assignment that we measure the distance of each home from the centrid. So distance from this centrid and this centrid and wherever which minimum assign it to that particular cluster and this process has been repeated again for both the centrid and once the distance has been measured on the improved or changed centrid again the assignment process has been started. So once you have moved and then measure the distance and then assignment also changes like it was done in the previous step and we continue this process till the time we have reached a location or a point where this change in assignment have stopped completely. So once we have reached this kind of place or this kind of scenario where as many time you measure the distance from the centrid to the different points your centroid does not change. This exercise or this point is called that your model has converged and at that point you can say okay these all group there is one group of these points or these homes. So this is one cluster and second one is this cluster. So this is how K means work. It has wide variety of applications. There is a function available in scikitlearn library. You can try implemented it. The intent of showcasing you this example of unsupervised learning was that we will be having two algorithms which come from unsupervised learning section of uh machine learning and these would be your restricted boltsman machines and autoenccoders which work on a similar methodology of unsupervised learning. So in a similar fashion like we started discussing in the beginning that where should be the location of these schools we can use a key means algorithm and initialize three points randomly and do this distance measure to each home and assign the homes to a cluster wherever the distance is minimum and we continue this process of measuring the distance and assigning it to the cluster till the time these value have been converged. The most important task for any data scientist is not to remember which library is required or what are the codes. In my understanding, the most important thing which data scientist should remember is that once you've been given a business problem, first you should be able to understand that what kind of problem it is. Whether it's a problem of supervised learning or it's an unsupervised learning. Given it's a supervised learning problem whether it falls into the regression type or a logistic regression type. If you can make these decisions then for implementing the algorithm you will find lot of help. In fact scikit learn would have initial codes for almost every algorithm. So you don't have to remember line of code and algorithms. All you should be able to do is once a problem has been given to you should be able to identify what kind of problem it is. Most of the time in unsupervised learning and specifically in C means kind of models we use this elbow method as indicator or help you understand that what is probably a number we should start with for starting the final implementation of your model. So let me give you an intuition how does it work? SSD stands for sum of squared errors and what it means is actually if I go back a little that suppose you have identified these two clusters. So sum of squared error would be that you take the centrid and measure the distance for the points which are associated with this cluster. So you measure the distance for each point in the orange group and sum square all the distances and the same exercise being done for the blue points and whatever the total number comes in after doing this exercise you will be getting what is the total number of sum of squared errors and if you have two clusters you would have some number and just for intuition I'm saying that this total sum of square is coming as 100 and that's only for intuition and example I'm taking this number to help you let's say there was one more cluster somebody identified here and all these three points though it's blue in color but I'm saying all these three point belong to this particular segment and rest of these points remained same as it was previously and as we saw with two clusters our sum of square was coming as 100 when we have three you can see that these points are bit far off from this particular cluster so if I'll be doing it with three clusters this distance would be a bit less given that now I have a point which is closer to these points And whatever error or distance these three points were adding it would be bit less given the cluster was here. And let's say this distance goes down to 95. I'm just making up some numbers. So probably what it is telling me that sum of squared is going down and probably I'm finding clusters which are closer or more closer to the actual data points. And as you would know that if I'll be increasing the number of clusters in the population, this distance would be going down hopefully and this distance can go up to zero when every point become a cluster itself. So if let's say I have 20 data points there and I assign that every point is a cluster in itself then just measuring the distance from the point which would be zero and overall SSD will become zero. So it may start from a very high number but it will be reducing with each cluster point or cluster you will be adding to your data. So this line which is sometime being called the elbow method what it's actually showing you. So if you had one cluster only anywhere in the population and you do some of the squared distances this was the distance. When you had two clusters this was the distance. When you had three these were the distance when you had four this was the distance. But when you had five the sum of squared error did not reduce much. So if you see it's like very less and after that even though you keep on adding different clusters the sum of squared errors is not going down. So as I was saying this process or this method is kind of indicative method and it gives you an intuition that if I have done it my cluster analysis with different number of clusters and I'm measuring the sum of squared errors for given number of clusters and I see that after four that the sum of squared error is not going down. It gives me an induction that probably I have found clusters which are more or less coherent and the population is not very much away from the centrid. From the point you can make it an assumption that probably four clusters is a good idea for my given population but as I also mentioned that it is just a indicative process. It's a good starting point but you need to see that how the distribution of your clusters look like whether they solve the business problem you're trying to solve or not and if not whether you need to further divide the clusters which your initial model has identified and here we'll be taking very quick introduction to a third type of learning which is called reinforcement learning. What it actually is we have seen from the two learning types the supervised learning and unsupervised learning. The first one was that we were trying to predict some dependent variable. In the second one, we are trying to identify some kind of structure in the data set or if I put it into other words that we are trying to identify some kind of coherent groups in the population. Third one is reinforcement learning and it's basically that an object or a system learns from the environment and there is no right or wrong answer given to the system explicitly or in the beginning itself like in the case of supervised learning. Here the object would be moving in the environment. Self-f flying helicopters where they fly on its own and they take the decision that what is the wind speed and what is the pressure around it and they correct their procedure accordingly and the objective they need to achieve is that they need to fly for a longer period of time. So here we are given an example. Let's say we have a robotic dog and somebody needs to train it to take correct decisions and correct things would be that it walking on the path where people needs to walk and it's not going down from the path and if some task is being given it's working correctly. So there are two components of reinforcement learning which is called reward and penalty. If the object or the system does the correct thing it receives some reward in terms of you know mathematical things. Obviously we will be providing everything in terms of mathematical numbers and if it does the wrong thing it receives a penalty and basis this thing it will keep on taking its decisions. So like a dog if it's walking correctly it receives the points like ball is being thrown if the robotic dogs go and pick it up it's a reward point. If it doesn't do the correct thing it receives a penalty. So most of like all these reinforcement learning agents working in a similar fashion. Some of you who are interested in implementing it, there is an algorithm called deep Q. It's an algorithm where you can design your own system and you can assign what are the rewards and penalty. Similar fashion reinforcement learning is also interacting with the space as I mentioned. So self-driving car is also one of the examples which would be receiving rewards and penalties based on whether it's running on the track, taking the right turns and moving at the correct speed, maintaining distance from other cars which are running. So reinforcement learning has a huge implementation or requirement for self-driving cars or some of the components of it not all some of the components also in self-driving car are supervised learning for the point that car needs to understand what the objects are in front of it and all other objects identification. So what are the real limitations of machine learning? Given that we already have all three type of algorithms supervised, unsupervised and reinforcement learning algorithm already available then why we want a new architecture or new type of algorithms for our artificial intelligence systems. First and foremost is the dimensions. And when I say dimension, it's like the type of data we get from lot of sources. Let's say we receive images which is gridlike image. So where are the pixels and what is the strength of pixels in the image natural language processing. So language data comes in a different length and you know the work is also different in the sense that suppose you need to design a machine learning algorithm which can do language translation and if you conceptualize this idea of language translation from a machine learning algorithm perspective your inputs become a sequence of words and your output is also sequence of word. And some of you who are working in machine learning algorithm try thinking that whether we have any algorithm currently available like logistic regression or decision tree which can help me even fit the algorithm or fit the problem. Leave aside how good the accuracy would be and all but these problems which come from a different type of data source and we trying to solve a different kind of problem like language translation or chatbot kind of problem where you give a sequence and it returns you a sequence. So these kind of architecture is already not available in machine learning. So that is one of the reasons that we need to identify some algorithms which can deal with such data sets like images and languages. And second, it can also fit different kind of models which are not only for predicting or classifying but also give you some kind of values like sequence I take an example of. So that is one of the reasons first we are looking for a different type of architecture for solving such problems. Then second problem which machine learning algorithms are not very good in dealing is the dimensionality. So we would have seen with a size of let's say thousand variables and let's say 100,000 rows and thousand columns probably you can still fit some of the machine learning algorithms on top of it. But given the kind of problems we are dealing with like images every image let's say it's uh 200 by 200 means 200 pixel by 200 pixel and it's a colored image it means there are three channels if you do the math 200 multiplied by 200 let's say this is your image and it's 200 by 200 because every image is kind of a matrix only and if it's a colored image actually colored image are being represented in system through three channels red green and blue so there would be three such grids but one top of the other. So number of pixels you need to have to represent your image in the system or in your algorithm would be 200x 200 by3 and then you calculate how many features it would be. If I if my math is correct it should be like 120,000 features. So even a simple image of such small dimensions you end up getting 120,000 features and plus if you're really working on a complex problem solving in terms of let's say an object identification in the images there may be five or six objects which you need to identify and you're dealing with let's say 100,000 images then your scale of data becomes so huge for any machine learning algorithm to easily handle it and your machine learning algorithms fail in terms of getting any interesting results out of it. So coming to solution part that we need an architecture which can not only read such data in terms of images but it is capable enough of dealing with such huge dimensions of data. So this is the second benefit which comes from the deep learning algorithms and we'll be discussing how do they manage such high dimensional data when we go and talk about different architectures. And third and the most important reason that we will be looking for a different kind of model structure or different kind of algorithm is for identifying the features. So in machine learning algorithms we as data scientists spend lot of time in kind of curating the important features. Either first you'll be scaling the features and after scaling you'll be creating the interaction variables then you'll be creating if the separator is not very clear then you need to introduce highdimensional data let's say it's your data point and if you see that line you're fitting is not separating clearly then some of you would be trying the higher order polomials of your input features. So all such things which not only difficult to you know come up there is lot of trial and error that which kind of transformation and which kind of variable creation. So what kind of variable will really work for classification problem that's first thing and second is if you're working on higher order polomials what is the correct order of polomial I need to create it and just to give you the scale of it let's say you are dealing with only 100 features and you need to create second order polomial with interaction of all these 100 features then you'll end up getting around 5,000 features from the second order polinomial only if you want to get third order polomial like cube variables or or the interaction of three variables together then these 100 variables will come around 170,000 features. So this creation of features is very very difficult. If we go and start creating these features on our own and our objective let's say to identify a television in the image and we have some pictures where we need to identify even though you have created those features manually and some of you who are working in the field of computer science for quite some time would know that earlier we used to use features like sift s if par features and hog features but these are like kind of static features for a given object but we may argue that this television is there in this picture here but in other picture it can be somewhere else. So the feature which I'm identifying it has to be spatial indifference that it can be anywhere in the image and same goes for language that if you're dealing with language data it should be not only able to understand the meaning of word or how does the word fit into the sentence but should also be able to understand that what is the context of each word but these word embeddings neural network help you understand that what are the related word to a given word and from that you make predictions. So these broad problems of machine learning algorithms one is they are not being able to play with or deal with different type of data like images and natural language. Second is the dimensional problem if the dimensionality goes in like 100,000 features and all. And third is this feature creation on its own. So these are the three basic reasons that one of you or all of you would be interested in going to one of the deep learning architecture for solving such problems. And fourthly, if I may add it that all the deep learning architectures given that we are putting lot of computational powers in them, they end up giving you a better accuracy both for classification and regression problems. So that that's the fourth benefit. And how does it really work? There are different stages in a deep learning and why they are called deep because it's not just input and output like we have seen in regression that you have a y and x is some kind of linear equation here we have different intermediate field like but there are lot of intermediary field for doing such complex calculations so that all these features which I mentioned that suppose you need to identify a television all such features get calculated at different stages one After the other and final stage you have very very refined features not only for image we are taking the image classification but any problem we are trying to solve through the multiple stages your model would be able to learn these intelligent features which are really important for your classification or regression or any such problem which we are trying to solve. So these were the few benefits for deep learning and these are actually the broad reasons that somebody would be interested in learning the deep learning algorithms. So this is the problem statement guys. We need to figure out if the bank nodes are real or fake and for that we'll be using artificial neural networks and obviously we need some sort of data in order to train our network. So let us see how the data set looks like. So over here I've taken a screenshot of the data set with few of the rows. In it data were extracted from images that were taken from genuine and forged banknotelike specimens. After that, wavelength transform tools were used to extract features from those images. And these are few features that I'm highlighting with my cursor. And the final column or the last column actually represents the label. So basically label tells us to which class that pattern represents whether that pattern represents a fake node or it represents a real node. Let us discuss these features and labels one by one. So the first feature or the first column is nothing but variance of a wavelength transformed image. The second column is about skewess. The third is courtesis of wavelength transformed image. And finally, fourth one is entropy of the image. After that when I talk about label which is nothing but my last column over here if the value is one that means the pattern represents a real load whereas when value is zero that means it represents a fake node. So guys let's move forward and we'll see what are the various steps involved in order to implement this use case. So over here we'll first begin by reading the data set that we have. We'll define features and labels. After that we are going to encode the dependent variable. And what is a dependent variable? It is nothing but your label. Then we are going to divide the data set into two parts. One for training, another for testing. After that, we'll use TensorFlow data structures for holding features, labels, etc. And TensorFlow is nothing but a Python library that is used in order to implement deep learning models or you can say neural networks. Then we'll write the code in order to implement the model. And once this is done, we will train our model on the training data. We'll calculate the error. The error is nothing but your difference between the model output and the actual output and we'll try to reduce this error and once this error becomes minimum we'll make prediction on the test data and we'll calculate the final accuracy. So guys let me quickly open my PyCharm and I'll show you how the output looks like. So this is my PyCharm guys. Over here I've already written the code in order to execute the use case. I'll go ahead and run this and I'll show you the output. So over here as you can see with every iteration the accuracy is increasing. So let me just stop it right here. All right. Till now any questions any doubts with respect to what is our use case? What is the data set about? So we'll move forward and we'll understand why we need neural networks. So in order to understand why we need neural networks, we are going to compare the approach before and after neural networks and we'll see what were the various problems that were there before neural networks. So earlier conventional computers use an algorithmic approach that is the computer follows a set of instructions in order to solve a problem and unless the specific steps that the computer needs to follow are known the computer cannot solve the problem. So obviously we need a person who actually knows how to solve that problem and he or she can provide the instructions to the computer as to how to solve that particular problem. Right? So we first should know the answer to that problem or we should know how to overcome that challenge or problem which is there in front of us. Then only we can provide instructions to the computer. So this restricts the problem solving capability of conventional computers to problems that we already understand and know how to solve. But what about those problems whose answer we have no clue of. So that's where our traditional approach was a failure. So that's why neural networks were introduced. Now let us see what was the scenario after neural networks. So neural networks basically process information in a similar way the human brain does. And these networks they actually learn from examples. You cannot program them to perform a specific task. They will learn from their examples from their experience. So you don't need to provide all the instructions to perform a specific task and your network will learn on its own with its own experience. All right. So this is what basically neural network does. So even if you don't know how to solve a problem, you can train your network in such a way that with experience it can actually learn how to solve the problem. So that was a major reason why neural networks came into existence. So these neural networks are basically inspired by neurons which are nothing but your brain cells and the exact working of the human brain is still a mystery though. So as I've told you earlier as well that neural networks work like human brain and so the name and similar to a newborn human baby as he or she learns from his or her experience we want a network to do that as well but we wanted to do very quickly. So here's a diagram of a neuron. Basically a biological neuron receives input from other sources combines them in some way perform a generally nonlinear operation on the result and then outputs the final result. So here if you notice these dendrites these dendrites will receive signals from the other neurons. Then what will happen? It'll transfer it to the cell body. The cell body will perform some function. It can be summation can be multiplication. So after performing that summation on the set of inputs via exxon it is transferred to the next neuron. Now let's understand what exactly are artificial neural networks. It is basically a computing system that is designed to simulate the way the human brain analyzes and process the information. Artificial neural networks has self-arning capabilities that enable it to produce better results as more data becomes available. So if you train your network on more data, it'll be more accurate. So these neural networks, they actually learn by example and you can configure your neural network for specific applications. It can be pattern recognition or it can be data classification, anything like that. All right. So because of neural networks, we see a lot of new technology has evolved from translating web pages to other languages to having a virtual assistant to order groceries online to conversing with chat bots. All of these things are possible because of neural networks. So in a nutshell, if I need to tell you artificial neural network is nothing but a network of various artificial neurons. All right. So let me show you the importance of neural network with two scenarios before and after neural network. So over here we have a machine and we have trained this machine on four types of dogs as you can see where I'm highlighting with my cursor and once the training is done we provide a random image to this particular machine which has a dog but this dog is not like the other dogs on which we have trained our system on. So without neural networks our machine cannot identify that dog in the picture as you can see it over here. Basically our machine will be confused. It cannot figure out where the dog is. Now when I talk about neural networks, even if we have not trained our machine on this specific dog, but still it can identify certain features of the dogs that we have trained on and it can match those features with the dog that is there in this particular image and it can identify that dog. So this happens all because of neural networks. So this is just an example to show you how important are neural networks. Now I know you all must be thinking how neural networks work. So for that we'll move forward and understand how it actually works. So over here I'll begin by first explaining a single artificial neuron that is called as perceptron. So this is an example of a perceptron. Over here we have multiple inputs x1 x2 dash till xn and we have corresponding weights as well. W1 for x1 w2 for x2 similarly wn for xn. Then what happens? We calculated the weighted sum of these inputs and after doing that we pass it through an activation function. This activation function is nothing but it provides a threshold value. So above that value my neuron will fire else it won't fire. So this is basically an artificial neuron. So when I talk about a neural network it involves a lot of these artificial neurons with their own activation function and their processing element. Now we'll move forward and we'll actually understand various modes of this perceptron or single artificial neuron. So there are two modes in a perceptron. One is training, another is using mode. In training mode, the neuron can be trained to fire for a particular input patterns. Which means that we'll actually train our neuron to fire on certain set of inputs and to not fire on the other set of inputs. That's what basically training mode is. When I talk about using mode, it means that when a taught input pattern is detected at the input, its associated output becomes the current output. Which means that once the training is done and we provide an input on which the neuron has been trained on. So it'll detect the input and will provide the associated output. So that's what basically using mode is. So first you need to train it then only you can use your perceptron or your uh network. So these were the two modes guys. Next up we'll understand what are the various activation functions available. So these are the three activation functions although there are many more but I've listed down three step function. So over here the moment your input is greater than this particular value your neuron will fire else it won't. Similarly for sigmoid and sign function as well. So these are three activation functions. There are many more that I've told you earlier as well. So these are the three majorly used activation functions. Next up what we are going to do we are going to understand how a neuron learns from its experience. So I'll give you a very good analogy in order to understand that and later on when we talk about a neural networks or you can say multiple neurons in a network I'll explain you the math behind it. I'll explain you the math behind learning how it actually happens. So right now I'll explain you with an analogy and guys trust me that analogy is pretty interesting. So I know all of you must have guessed it. So these are two beer mugs and all of you who love beer can actually relate to this analogy a lot. And I know most of you actually love beer. So that's why I've chosen this particular analogy so that all of you can relate to it. All right, jokes apart. So fine guys, so there's a beer festival happening near your house and you want to badly go there. But your decision actually depends on three factors. First is how is the weather, whether it is good or bad. Second is your wife or husband is going with you or not. And the third one is any public transport is available. So on these three factors, your decision will depend whether you will go or not. So we'll consider these three factors as inputs to our perceptron and we'll consider our decision of going or not going to the beer festival as our output. So let us move forward with that. So we'll move forward and we'll see what are the various inputs that I'm talking about. So the first input is how is the weather? We'll consider it as x1. So when weather is good it'll be one and when it is bad it'll be zero. Similarly your wife is going with you or not. So that be your x2. If she is going then it's one. If she's not going then it's zero. Similarly for public transport if it is available then it is one else it is zero. So these are the three inputs that I'm talking about. Let's see the output. So output will be one when you're going to the beer festival and output will be zero when you want to relax at home. You want to have beer at home only. You don't want to go outside. So these are the two outputs whether you are going or you're not going. Now what a human brain does over here. Okay fine I need to go to the beer festival but there are three things that I need to consider. But will I give importance to all these factors equally? Definitely not. There'll be certain factors which will be of higher priority for me. I'll focus on those factors more. Whereas few factors won't affect that much to me. All right. So let's prioritize our inputs or factors. So here our most important factor is weather. So if weather is good, I love beer so much that I don't care even if my wife is going with me or not or if there is a public transport available. So I love beer that much that if weather is good that definitely I'm going there. That means when x1 is high output will be definitely high. So how we do that? How we actually prioritize our factors or how we actually give importance more to a particular input and less to another input in a perceptron or in a neuron. So we do that by using weights. So we assign high weights to the more important factors or more important inputs and we assign low weights to those particular inputs which are not that important for us. So let's assign weights guys. So weight w is associated with input x1, w2 with x2 and similarly w3 with x3. Now as I've told you earlier as well that weather is a very important factor. So I'll assign a pretty high weight to weather and I'll keep it as six. Similarly w2 and w3 are not that important. So I'll keep it as 22. After that I've defined a threshold value as five which means that when the weighted sum of my input is greater than five then only my neuron will fire or you can say then only I'll be going to the b festival. All right. So I'll use my pen and we'll see what happens when weather is good. So when weather is good, our x1 is 1. Our weight is six. We'll multiply it with six. Then if my wife decides that she is going to stay at home and she will probably be busy with cooking and she doesn't want to drink beer with me, so she's not coming. So that input becomes zero. 0 into 2 will actually make no difference because it'll be zero only. Then again there is no public transport available also. Then also this will be 0 into 2. So what output I get here? I get here as six. And notice the threshold value it is five. So definitely six is greater than five. That means my output will be one or you can say my neuron will fire or I'll actually go to the beer festival. So even if these two inputs are zero for me that means my wife is not willing to go with me and there is no public transport available but weather is good which has very high weight value and it actually matters a lot to me. So if that is high it doesn't really matter whether the two inputs are high or not I will definitely go to the BF festival. All right now I'll explain you a different scenario. So over here our threshold was five but what if I change this threshold to three. So in that scenario even if my weather is not good uh I'll give it the zero. So 0 into 6 but my wife and public transport both are available. All right. So 1 into 2 + 1 into 2 which is equal to 4 and it is definitely greater than three. Then also my output will be one. that means I will definitely go to the beer festival even if weather is bad and my neuron will fire. So these are the two scenarios that I have discussed with you. All right. So there can be many other ways in which you can actually assign weight to your uh problem or to your learning algorithm. So these are the two ways in which you can assign weights and prioritize your inputs or factors on which your output will depend. So obviously in real life all the inputs or all the factors are not as important for you. So you actually prioritize it and how you do that in perceptron you provide high weight to it. This is just an analogy so that you can relate to a perceptron to a real life. We'll actually discuss the math behind it later in the session as to how a network or a neuron learns. All right. So how the weights are actually updated and how the output is changing that all those things we'll be discussing later in the session. But my aim is to make you understand that you can actually relate to a real life problem with that of a perceptron. All right? And in real life problems are not that easy. They are very very complex problems that we actually face. So in order to solve those problems a single neuron is definitely not enough. So we need networks of neuron and that's where artificial neural network or you can say multi-layer perceptron comes into the picture. Now let us discuss that multi-layer perceptron or artificial neural network. So this is how an artificial neural network actually looks like. So over here we have multiple neurons in present in different layers. The first layer is always your input layer. This is where you actually feed in all of your inputs. Then we have the first hidden layer. Then we have second hidden layer and then we have the output layer. Although the number of hidden layers depend on your application on what are you working what is your problem. So that actually determines how many hidden layers you'll have. So let me explain you what is actually happening here. So you provide in some input to the first layer which is nothing but your input layer. You provide inputs to these neurons. All right? And after some function the output of these neurons will become the input to the next layer which is nothing but your hidden layer one. Then these hidden layers also have various neurons. These neurons will have different activation functions. So they'll perform their own function on the inputs that it receives from the previous layer. And then the output of this layer will be the input to the next hidden layer which is hidden layer 2. Similarly, the output of this hidden layer will be the input to the output layer and finally we get the output. So this is how basically an artificial neural network looks like. Now let me explain you this with an example. So over here I'll take an example of image recognition using neural networks. So over here what happens? We feed in a lot of images to our input layer. Now this input layer will actually detect the patterns of local contrast and then we'll feed that to the next layer which is hidden layer 1. So in this hidden layer one the face features will be recognized we'll recognize eyes nose ears things like that and then that will be again fed as input to the next hidden layer and in this hidden layer we'll assemble those features and we'll try to make a face and then we'll get the output that is the face will be recognized properly. So if you notice here with every layer we are trying to get a more abstract version or the generalized version of the input. So this is now basically an artificial neural network how it works. All right. And there's a lot of training and learning which is involved that I'll show you now training a neural network. So how we actually train our neural network. So basically the most common algorithm for training a network is called back propagation. So what happens in back propagation after the weighted sum of inputs and passing through an activation function and getting the output. We compare that output to the actual output that we already know. We figure out how much is the difference. We calculate the error and based on that error what we do we propagate backwards and we'll see what happens when we change the weight will the error decrease or will it increase and if it increases when it increases by increasing the value of the variables or by decreasing the value of variables. So we kind of calculate all those things and we update our variables in such a way that our error becomes minimum and it takes a lot of iterations. Trust me guys it takes a lot of iterations. We get output a lot of times and then we compare it with the model with the actual output. Then again we propagate backwards. We change the variables. Then again we calculate the output. We compare it again with the desired output or the actual output. Then again we propagate backwards. So this process keeps on repeating until we get the minimum value. All right. So there's an example that is there in front of your screen. Don't be scared of the terms that I used. I'll actually explain you with an example. So this is the example over here. We have 0 1 and two as inputs. And our desired output or the output that we already know is 0 1 and four. All right. So over here we can actually figure out that desired output is nothing but twice of your input. But I'm training a computer to do that. Right? The computer is not a human. So what happens? I actually initialize my weight. I keep the value as three. So the model output will be 3 into 0 is 0. 3 into 1 is 3. 3 into 2 is 6. Now obviously it is not equal to your desired output. So we check the error. Now the error that we have got here is 0 1 and 2 which is nothing but your difference. So 0 - 0 is 0 3 - 2 is 1 6 - 4 is 2. Now this is called an absolute error. After squaring this error we get square error which is nothing but 0 1 and 4. All right. So now what we need to do we need to update the variables. We have seen that the output that we got is actually different from the desired output. So we need to update the value of the weight. So instead of three our computer makes it as four. After making the value as four, we get the model output as 0 4 and 8. And then we saw that the error has actually increased. Instead of decreasing, the error has increased. So after updating the variable, the error has increased. So you can see that square error is now 0 4 and 16. And earlier it was 0 1 and 4. That means we cannot increase the weight value right now. But if we decrease that, make it as two, we get the output which is actually equal to desired out. But is it always the case that we need to only decrease the weight? Definitely not. So in this particular scenario, whenever I'm increasing the weight, error is increasing and when I'm decreasing the weight, error is decreasing. But as I've told you earlier as well, this is not the case every time. Sometimes you need to increase the weight as well. So how we determine that? All right, fine guys. This is how basically a computer decide whether it has to increase weight or decrease away. So what happens here? This is a graph of square error versus weight. So over here what happens? Suppose your square error is somewhere here and your computer it starts increasing the weight in order to reduce the square error and it notices that whenever it increases the weight square error is actually decreasing. So it'll keep on increasing until the square error reaches a minimum value and after that when it tries to still increase the weight the square error will increase. So at that time our network will recognize that whenever it is increasing the weight after this point error is increasing. So therefore it will stop right there and that will be our weight value. Similarly there can be one more scenario. Suppose if we increase the weight but then also the square error is increasing. So at that time we cannot increase the weight. At that time computer will realize okay fine whenever I'm increasing the weight the square error is increasing. So it'll go in the opposite direction. So it'll start decreasing the weight and it'll keep on doing that until the square error becomes minimum. And the moment it decreases more the square error again increases. So our network will know that whenever it decreases the weight value the square error is increasing. So that point will be our final weight value. So guys this is what basically back propagation in a nutshell is. If you have any questions or doubts you can go ahead and ask me. All right fine we have no doubts here. Fine. So we'll move forward and now is the correct time to understand how to implement the use case that I was talking about at the beginning. That is how to determine whether a node is fake or real. So for that I'll open my PyCharm. This is my PyCharm again guys. Uh let me just close this. All right. So this is the code that I've written in order to implement the use case. So over here what we do we import the first important libraries which are required. Mattplot lab is used for visualization. TensorFlow we know in order to implement the neural networks. Numpy for arrays. Pandas for reading the data set. Similarly sklearn for label encoding as well as for shuffing and also to split the data set into training and testing task. All right fine guys. So we'll begin by first reading the data set as I've told you earlier as well when I was explaining the steps. So what I'll do I'll use pandas in order to read the CSV file which has the data set. After that I'll define features and labels. So x will be my feature and y will contain my label. So basically x includes all the columns apart from the last column which is the fifth one. And because the indexing starts from zero that's why we have written 0 till fourth. So it won't include the fourth column. All right. And so our last column will actually be our label. Then what we need to do, we need to encode the dependent variable. So the dependent variable as I've told you earlier as well is nothing but your label. So I've discussed encoding in TensorFlow tutorial. You can go through it and you can actually get to know why and how we do that. Then what we have done, we have read the data set. Then what we need to do is to split our data set into training and testing. And these are all optional steps. You can print the shape of your training and test data. If you don't want to do it yourself, fine. Then we have defined learning rate. So learning rate is actually the steps in which the weights will be updated. All right. So that is what basically learning rate is. Then when we talk about epoch means iterations. Then we have defined cause history that will be an empty numpy array and it shape will be one and it'll include the flow type objects. Then we have defined end which is nothing but your x shape of axis one which means your column. Then we'll print that. After that we have defined the number of classes. So there can be only two class whether the node can be fake or it can be real. And this model path I've given in order to save my model. So I've just given a path where I need to save it. So I'll just save it here only in the current working directory. Now is the time to actually define our neural network. So we'll first make sure that we have defined the important parameters like hidden layers, number of neurons in hidden layers. So I'll take 10 neurons in every hidden layer and I'm taking four layers like that. Then x will be my placeholder and the shape of this particular placeholder is none, n dim. N dim value I'll get it from here and none can be add any value. I'll define one variable w and I'll initialize it with zeros and this will be the shape of my weight. Similarly for bias as well this will be the particular shape and there will be one more placeholder ydash which will actually be used in order to provide us with the actual output of the model. There'll be one model output and there'll be one actual output which we use in order to calculate the difference. Right? So we'll feed in the actual values of the labels in this particular placeholder ydash. And now we'll define the model. So over here we have name the function as multi-layer perceptron. And in it we'll first define the first layer. So the first hidden layer and we are going to name it as layer_1 which will be nothing but the matrix multiplication of x and weights of h1 that is the hidden layer 1 and that'll be added to your biases b1. After that we'll pass it through a sigmoid activation function. Similarly in layer 2 as well matrix multiplication of layer 1 and weights of h2. So if you can notice layer 1 was the network layer just before the layer 2 right. So the output of this layer one will become input to the layer two and that's why we have written layer_1 it'll be multiplied by weight h2 and then we'll add it with the bias. Similarly for this particular hidden layer as well and this particular layer as well but over here we are going to use the rail activation function instead of sigmoid. Then we are going to define the weights and biases. So this is how we basically define weights. This is how we basically define weights. The weights h1 will be a variable which will be a truncated normal with the shape of n dim and nc hidden_1. So these are nothing but your shapes. All right. And after that what we have done we have defined biases as well. Then we need to initialize all the variables. So all these things actually I've discussed in brief when I was talking about tensorflow. So you can go through tensorflow tutorial at any point of time if you have any question. We have discussed everything there. Since in tensorflow we need to initialize the variables before we use it. So that's how we do it. We first initialize it and then we need to run it. That's when your variables will be initialized. After that we are going to create a saver object and then finally I'm going to call my model. And then comes the part where the training happens. Cost function. Cost function is nothing but you can say an error that will be calculated between the actual output and the model output. All right. So y is nothing but our model output and ydash is nothing but actual output or the output that we already know. All right. And then we are going to use a gradient descent optimizer to reduce error. Then we are going to create a session object as well. And finally what we are going to do we are going to run the session. So this is how we basically do that. For every epoch we will be calculating the change in the error as well as the accuracy that comes after every epoch on the training data. After we have calculated the accuracy on the training data, we going to plot it for every epoch how the accuracy is. And after plotting that we going to print the final accuracy which will be on our test data. So using the same model we'll make prediction on the test data. And after that we are going to print the final accuracy and the mean squared error. So let's go ahead and execute this guys. All right. So training is done and this is the graph we have got for accuracy versus epoch. This is accuracy. Y-axis represents accuracy whereas this is epoch. We have taken 100 epochs and our accuracy has reached somewhere around 99%. So with every epoch it is actually increasing apart from a couple of instances it is actually keep on increasing. So the more data you train your model on it'll be more accurate. Let me just close it. So now the model has also been saved where I wanted it to be. This is my final test accuracy and this is the mean squared error. All right. So these are the files that will appear once you save your model. These are the four files that I've highlighted. Now what we need to do is restore this particular model and I've explained this in detail how to restore a model that you have already saved. So over here what I'll do I'll take some random range. I've taken it actually from 754 to 768. So all the values in the row of 754 and 768 will be fed to our model and our model will make prediction on that. So let us go ahead and run this. So when I'm restoring my model, it seems that my model is 100% accurate for the values that I have fed in. So whatever values that I have actually given as input to my model, it has correctly identified its class whether it's a fake node or a real node because zero stands for fake node and one stands for real node. Okay. So original class is nothing but which is there in my data set. So it is zero already and what prediction my model has made is zero. That means it is fake. So accuracy becomes 100%. Similarly for other values as well. Fine guys. So this is how we basically implement the use case that we saw in the beginning. So in the slide you can notice that I've listed out only two applications although there are many more. So neural networks in medicine. Artificial neural networks are currently a very hot research area in medicine and it is believed that they will receive extensive application to biomedical systems in the next few years and currently the research is mostly on modeling parts of human body and recognizing diseases from various scans for example it can be cardiograms cat scans ultrasonic scans etc. And currently the research is going mostly on two major areas first is modeling and diagnosing the cardiovascular system. So neural networks are used experimentally to model the human cardiovascular system. Diagnosis can be achieved by building a model of the cardiovascular system of an individual and comparing it with the real-time physiological measurements taken from the patient. And trust me guys, if this routine is carried out regularly, potential harmful medical conditions can be detected at an early stage and thus make the process of combating disease much easier. Apart from that, it is currently being used in electronic noses as well. Electronic noses have several potential applications in tele medicine. Now, let me just give you an introduction to tele medicine. Tele medicine is a practice of medicine over long distance via a communication link. So, what the electronic noses will do, they would identify odors in the remote surgical environment. These identified odors would then be electronically transmitted to another site where an door generation system would recreate them. Because the sense of the smell can be an important sense to the surgeon. Teley smell would enhance telepresent surgery. So these are the two ways in which you can use it in medicine. You can use it in business as well guys. So business is basically a diverted field with several general areas of specialization such as accounting or financial analysis. Almost any neural network application would fit into one business area or financial analysis. Now there is some potential for using neural networks for business purposes including resource allocation and scheduling. I've listed down two major areas where it can be used. One is marketing. So there is a marketing application which has been integrated with a neural network system. The airline marketing tactician is a computer system made of various intelligent technologies including expert systems. A feed forward neural network is integrated with the AMT which is nothing but airline marketing tactician and was trained using back propagation to assist the marketing control of airline seat location. So it has wide applications in marketing as well. Now the second area is credit evaluation. Now I'll give you an example here. The HNC company has developed several neural network applications and one of them is a credit scoring system which increases the profitability of existing model up to 27%. So these are few applications that I'm telling you guys neural network is actually the future. People are talking about neural networks everywhere and especially after the introduction of GPUs and the amount of data that we have now neural network is actually spreading like plague right now. So what is a transformer? Transformers operate on a concept called sequencetosequence learning. Essentially they take a sequence of tokens as an input and predict the next token in the output. A great example of this is language translation. Imagine inputting good morning in English and the transformer process this and outputs the translation in languages like Japanese, Korean or German. The key is how it efficiently process the relationship between words. Since we know what a transformer is, let's dig a bit deep about them. A transformer has two primary components encoder and decoder. Encoder identifies relationship between parts of the input sequence whereas the decoder uses these relationships to generate the output sequence. This division is what allows transformers to handle task like text translation or summarizations with remarkable accuracy. Now that we have the idea of transformers, let's discuss how they evolve. Before transformers, there were other neural networks like RNN, recurrent neural networks invented by David Rumlhur in 1986. However, RNN's face significant challenges. They would forget early parts of the sequence as they processed longer ones and couldn't handle dependencies efficiently. Additionally, RNN relied on recurrence which made them inefficient and incapable of parallelization. Then came long short-term memory introduced by Howeter and Smith Huber in 1997. Long short-term memory improved by remembering sequences for a longer duration and addressing some of the memory issues in RNN. However, they were slow to train and difficult to manage at scale. Finally, transformers transformed neural networks. First introduced in the landmark paper, attention is all you need. Transformers addressed all the problems faced by RNN and LSTMs. They used a completely attention-based mechanism, eliminating reliance on recurrence. This made transformers capable of remembering context efficiently, training faster and being parallelized, enabling multitasking and significantly speeding up processes. Now, let's discuss on the attention mechanism. Think about the sentence. This cat wants to jump on the box. The attention mechanism identifies the most relevant parts of this sentence like cat, jump and box and focuses on these elements while processing the data. Now that we know how transformers have evolved, now let's discuss their architecture. A transformer consists of two main components, an encoder and a decoder. Each typically consisting six layers. Inside the encoder, there is one attention layer and one feed forward layer. While the decoder contains two attention layers and one feed forward layer. The magic of parallelism comes from how data is fed into the network. In the attention layer, all the words are processed simultaneously with each word forming combinations with other in the sentence. This allows the model to capture relationships and context efficiently. After processing in the attention layer, the data is sent to the feed forward layer where it is learned layer by layer. The input to the encoder and decoder are the raw input embeddings which are numerical representation of words. On top of these embeddings, positional encodings are added to help the model understand the position and the order of each word in the sequence. If we simplify embeddings, there are essentially vector representation of words in an n dimensional space. At the top of the architecture there are two layers of the output probabilities converting the final output into the form of human can understand. These inputs are represented as vectors with their length corresponding to the size of vocabulary. Now what truly makes transformer unique is the inclusion of normalization layers which normalize the output from sub layers. Additionally, skip connections, the dark arrows in the architecture, forward critical information that bypasses self attention or fed forward layers directly to the normalization layers. This ensures the model does not forget important details and effectively passes vital information further into the network. Now, moving forward, let's discuss why transformers are important. Transformers are vital because they utilize semi-supervised learning. They are trained on massive unlabelled data sets enabling them to generalize across a wide range of task. Unlike older models, transformers don't need to process data sequentially. Their attention mechanism allow them to focus on the most relevant context which significantly speeds up training. Transformers revolutionize data processing by eliminating the need to handle data sequentially allowing for parallel processing and significantly enhancing efficiency. The attention mechanism lies at the core of transformers, enabling the model to focus on the most relevant parts of the input sequence and improving accuracy and understanding of context. Furthermore, transformers excel at providing context, ensuring that the meaning of each word or token is accurately interpreted within its surroundings. Lastly, these models dramatically speeds up the training process, making them faster and more efficient compared to traditional neural networks, thus redefining AI's capabilities across diverse applications. Now that we know why transformers are important, let's discuss some applications. We have OpenAI's GPT, a groundbreaking model that leverages the power of transformers for natural language processing task. Additionally, Google has developed several transformer-based models including vision transformer for image recognition but birectional encoder representations from transformers for understanding the context of words in a sentence and T5 which stands for texttoext transfer transformer for a wide range of text generation task. Microsoft has also contributed with debt decoding enhanced bird with disentangled attention a model designed to improve contextually understanding and enhance NLP applications. These models demonstrate the versatility and impact of transformer architecture across various domains. Now that we know the application of transformers, how about checking their realtime products? Transformers have become an integral part of many real world products that we use style. Examples include Grammarly which leverages transformers for advanced grammar and writing assistance. Google search and its translation tools powered by models like bird and t5 and chat GPD open AI's conversational AI that relies on the generative pre-train transformer architecture. Additionally, Meta's deep fake detector uses transformer-based models for facial recognition task. These applications highlight how transformers have revolutionized technology, seamlessly integrating it into tools that enhance our everyday lives. In conclusion, transformers are changing the tech world by enabling smarter, faster, and more efficient AI systems. Whether it's generating text, translating languages, or enhancing search engines, these models are the cornerstone of modern AI. So what are RNN's, right? Well, RNN basically stand for recurrent neural network and we usually use this in order to deal with a sequential data. Sequential data can be something like a time series data or a textual data of any format. So why should one use RN and write? But this is because there's a concept of internal memory here. RNN can remember important things about the input it has received which allows them to be very precise in predicting what can be the next outcome. So this is the reason why they are performed or preferred on a sequential data algorithm. Okay. And some of the examples of sequential data can be something like time series, speech, text, financial data, audio, video, weather and many more. Although RNN were the state-of-the-art algorithm for dealing with sequential data, they come up with their own drawbacks. And some of the popular drawbacks over here can be like due to the complication or the complexity of the algorithm the neural network is pretty slow to train and as there are huge amount of dimensions here the training is very long and difficult to do. Okay. Apart from that the most decisive feature for RNN or for the improvement in RNN is that of a vanishing gradient. What this vanishing gradient is is that you know when we go deeper and deeper into our neural network the previous data is lost. This is because of a concept called as vanishing gradient and due to this we cannot work on a large or a longer sequence of data. Okay. To overcome this we came up with some new or upgrades to the current recurrent neural networks or RNN. Starting off with birectional recurrent neural network. You see birectional recurrent neural network connect two hidden layers of opposite direction into the same output. With this form of generative deep learning, the output layer can get information from past future states simultaneously. So as you can see here we have two layers over here and as they are birectional what happens is when the algorithm feels that it is kind of losing its gradients or the previous data, it can go back and get the data from the past. So why do we need birectional recurrent neural network? Well, birectional recurrent neural network duplicates RNN processing chain so that the input process both forward and reverse time order thus allowing birectional recurrent neural network to look into future context as well. The next one is long short-term memory. Long short-term memory or also sometime referred to as LSTM is an artificial recurrent neural network architecture used in the field of deep learning. Unlike standard feed forward neural network, LSTM has a feedback connections. It can not only process single data point but also the entire sequence of data. So as you can see here from what I'm trying to say is with LSTM or longsh short-term memory it has something like you know we can feed a longer sequence compared to what it was with birectional RNN or RNN. So why is LSTM better than RNN? We can say that when we move from RNN to LSTM we are introducing more and more control over the sequence of the data that we can provide. whereas LSTM gives us more controllability and there's better results. All right. So the next type of recurrent neural network is the gated recurrent neural network or also referred to as GRUs. You see GRU is a type of recurrent neural network that is in certain cases is advantageous over long short-term memory. GRU makes use of less memory and also is faster than LSTM. But thing is LSTMs are more accurate while using longer data sets. I'm sure by now you might have got a hint about the trend that has lead to the improvement. Right? So the trend over here is you know the model should be capable of remembering and taking in on a longer input sequence. The game changer part for the sequential data was developed when we came up with something called as transformers. And this paper was something which is based on a concept called as attention is everything. All right. So let's take a look at this. The paper attention is all you need introduces a novel architecture called as transformers. Like LSCM transformers is an architecture for transforming one sequence into another while helping other two parts that is encoders and decoders. But it differs from previously described sequence to sequence model because it does not work like GRUs. Okay. So it does not implements recurrent neural networks. Recurrent neural network until now were one of the best ways to capture the timely dependence on a sequence. However, the team presenting this paper that is attention is all you need prove that an architecture with only attention mechanism does not use RNN can improve its result in translation task and other NLP task. One of the best examples for transformers is Google's bird. So what exactly is this transformer? Right, we see here we have encoder on the top and decoder on the bottom. Both encoder and decoder are comprised of modules that can stick onto the top of each other multiple times. So what happens here is the inputs and outputs are first embedded into n dimension space since we cannot use this directly. So we obviously have to encode our inputs whatever we are providing here. One slight but important part of this model is the positional encoding of different words. Since we have no recurrent neural network that can remember how sequence are fed into the model. we need to somehow give every word or part of our sequence a relative position since a sequence depends on the order of the elements. Okay, these positions are added to the embedded representation of each words. All right, so this was the brief about transformers. So let us now move ahead and see some of the popular language models that are available in the market. All right, so let us now start off by understanding OpenAI's GPT3. The successor to GPT and GPT2 is the GPT3 and is one of the most controversial pre-trend models by OpenAI. The large scale transformer-based language model has been trained on 175 billion parameters which is 10 times more than any previous non-sparse language model. The model has been trained to achieve strong performance on many NLP data set including task like translation, answering questions as well as several other tasks. Then we have Google's bird. Bird stands for birectional encoder representations from transformers. A is a pre-trained NLP model which is developed by Google in 2018. With this anyone in the world can train either their own question answering module with up to 30 minutes on a single cloud TPU or few hours using single GPU. The company then released this showcasing the performance of 11 NLP task including very competitive Stanford data set questions. Unlike other language model, bird has only been pre-trained on 250 million words of Wikipedia and 800 million words of book corpus and has been successfully used as a pre-trained model in deep neural network. According to researchers, bird has achieved 93% accuracy which has surpassed any previous language models. Next, we have Elmo. Elmo also known as embedding for language model is a deep contextualized word representation that models syntax and semantic words as well as their logistic context. The model developed by Alan LLP has been pre-trained on a huge text corpus and learn functions from birectional models that is by LM. Elmo can easily be added to their existing models which drastically improves the features of functions across vast NLP problem including answering questions, textual sentiment and sentiment analysis. Now let's answer the fundamental question that is what exactly is generative AI? Generative AI refers to algorithm capable of creating new content with a text, images, audio, or even videos. It's like having a creative AI assistant that can take a simple input and produce engaging outputs. For example, GPT and Llama can write essays or code while image generation models like DAL E and stable diffusion can visualize unique scenes from descriptions. But let's look at some of the popular tools driving this innovation. Well, some of the standard tools in generative AI includes GitHub copilot which assists developers with code suggestions and charg for text based interactions. Image generation tools like stable diffusion and midjourney helps creators bring visual concepts to life. Google's Gemini merges text and image capabilities while Adobe Firefly extends AI's reach to creative suits. So if you want to know how to use these tools, then check out our Generative AI examples video link in the description. So you might wonder where are these tools being applied. Now let's explore them. Generative AI is transforming multiple creative fields. Image generation tools power visual design. Music compositions algorithm create original scores and AI assist video editors in automatic task. LLM help generate and translate text while code generation tools like GitHub copilot boost developer efficiency. AI generated voices are even being used in audio books and voice assistance. So now let's take some of these tools and check. So this time we will use Ptory AI and Flicky AI. First let's explore Ptory AI. So for that let's go to its site and check its functions. So we are at the Ptory AI site. And on the left side we have the home project and brand kits. And on the main screen we have different features Ptory AI provides. So let's choose text to video. Here let's write some names and description and press generate. Pictory AI is a tool designed for video creators that helps transform long- form content such as articles or blog post into short engaging videos. It uses AI to automatically extract key highlights and create professionallook videos with minimal efforts. Due to its simplicity and time-saving capabilities, Pictory AI is a popular for social media content creation and marketing. Now let's see our next tool which is Flicky AI. So now we are at the flicky.ai site. Here we have different features like videos where you can create videos from all of these blogs, prompts etc. You can also create audios from these features and then we also have a design feature and on the left hand side you can see options like files, templates, brand kits, voice clones etc. So now let's take an idea and convert it into a video. Now let's write our topic and generate. Flicky AI is a content creation tool that turns text into videos using AI generated voices and visuals. It helps users create professional videos quickly by pairing written content with the stock images, animations and voiceovers. Flicky is ideal for marketers, content creators, and educators looking to create engaging video content efficiently. Now that we have seen the applications, so let's step back and look at the journey that brought us here. So basically our journey starts in 1947 with Alan Turing's concept of intelligent machines. By 1961, Joseph Venbomb introduced ELA, the first chatbot. The 1980s saw the birth of recurrent neural networks, while 1997 brought longshortterm memory networks to tackle sequential data. And then GANs emerged in 2014 transforming creative task. Fast forward to 2017 when transformers like GPT entered the scene. By 2023, GPT 3.5 and Google's palm marked significant milestones. And by 2025, we are on the brink of AI breakthroughs in chemistry and genome editing. So what exactly are these LLMs and why are they so powerful? An LLM or large language models analyzes and understands natural language using machine learning. Examples include OpenAI's GPT, Google Spetas, Llama. These models drive applications such as chatbots, language translation, and more by learning from extensive data to predict and generate text sequences. But before this, there was a very famous term called language model. A language model is a machine learning model that uses probability statistics and mathematics to predict the next sequence of words. Suppose you have a sentence like I have a boy who is my dash. Here if we ask a language model to predict the next word, it considers the context provided by the words before the blank. Based on common usage patterns from its training data, it may predict words like boyfriend, brother, or friend which fit naturally. However, it's less likely to predict colleague or sibling as those words may not commonly follow these type of phrases. So, this process shows how language models predict text by calculating probabilities for each possible word based on their likelihood in context. So, when a language model is trained on massive amounts of diverse text, it gains a wider vocabulary and more understanding of language, enabling it to make more accurate predictions. For example, if we give it a phrase like you are a dash to me, a model trained on extensive data might suggest various fitting words for example friend, inspiration or anything else. So based on the sentiment or context it has learned from the data. Now here reinforcement learning is used to improve the model's responses over time. By giving feedbacks be it positive or negative on the responses we help the model learn which type of responses are preferred in specific context. For example, if the model frequently misinterprets the tone or intent, the reinforcement learning helps adjust its productions to be more contextually appropriate and aligned with the intended meaning. But what do these models look like under the hood? Well, LLMs are built on neural networks composed of input, hidden, and output layers. The hidden layer process information to learn complex patterns and more layers means the model can capture deeper insights. This structure allows LLM to perform task from generating text to complex code completions. Now, how do these layers interact and function in real time? Now, LLM is based on the transformer and a transformer uses deep learning to process any information coming to it. Now, let me tell you a story of three friends. Imagine we have three characters. First is our friend. The next character is Minion Bob and the third character is Gru. So, our friend asks Bob, "What's the price of the jet? It must be $50,000." Minion Bob isn't sure. So he goes to Gru and asks, "Is the jet $50,000?" Grrew replies, "No, it's $70,000." In this back and forth, minion Bob is like the neural network layer trying to make an accurate guess. So each time he goes back to group, like receiving more data or feedback, he gets corrected if his guess is wrong, leading him to refine his response. Now, after the first check, Min and Bob returns to our friend saying, "I guess it's more than $60,000." Our friend assumes it might be around $65,000. And sends Bob back to Gru to verify. Again, Gru corrects him. No, it's actually $70,000. So, this process repeats with Bob adjusting his guess each time. Eventually, he learns that the correct answer is $70,000 and updates his knowledge. So just like Minion Bob, neural networks make initial guesses based on available information with each feedback loop like Bob going back to group, the model's hidden layers adjust the parameters to refine its guesses, ultimately arriving at the most accurate prediction possible. So after getting corrected multiple times, Minion Bob's guesses improve until he knows the price is $70,000. Similarly, in a neural network, gradually learning the correct answer through training. So once the network learns it can give accurate answers in future cases without checking every time. Now let us move on to understand how LLMs work. LLMs begin the collection of data sets then tokenize text and break it into a manageable pieces. Using a transformer architecture they process the data sequence all at once leveraging vast training data. LLMs contain millions of learned parameters that predict the text tokens and generate coherent outputs. Models often undergo pre-training for general knowledge and fine-tuning for specific task. So now let's see some practical uses of LLMs. LLMs power content generation creating anything from articles to code. They excel in language translation, enhanced search engines, personalized recommendation, code development assistance and sentiment analysis which also owe much to LLM's predictive capabilities. So guys, are you ready to use all that knowledge in coding and witness how these LLMs come together to drive innovation? Whether through developing applications, analyzing data or building smart assistance, the gear of technology keep turning to unlock AI's full potential. So now let us look at our problem statement. So one of the difficulties in the healthcare industry is effectively evaluating medical pictures such as MRIs, CT scans and X-rays in order to identify anomalies and illnesses. This procedure takes a lot of time and calls from specialized understanding. Automated methods must be developed to help medical personnel recognize possible health problems in medical imaging. In order to provide better patient care, a system that integrates cutting edge machine learning models with image analysis can greatly help in the early detection of diseases including cancer, infections, and other illnesses. So, the method uses generative AI to evaluate medical photos and generate a thorough diagnosis report based on the findings. This technology allows users to upload medical images which the AI model then processes. Now let us build our project on a medical image analysis application using streamlit Python and an LLM of Google Gemini AI
Original Description
🔥Agentic AI Training Course - Master AI Agents: https://www.edureka.co/agentic-ai-training-course
🔥Integrated MS+PGP Program in Data Science & AI:https://www.edureka.co/dual-certification-programs/ms-data-science-pgp-gen-ai-ml-birchwood
Explore the future of artificial intelligence with Agentic AI! In this video, we dive into the exciting developments and advancements that are set to shape the industry in 2026. We will learn what Agentic AI is and how it goes beyond traditional AI by acting with purpose and autonomy. You’ll discover 5 powerful secrets behind Agentic AI—from multi-agent collaboration to real-world tool integration and ethical guardrails. By the end, you’ll understand how Agentic AI is transforming industries and why it’s the future of intelligent systems.
Join us as we discuss the latest trends, innovations, and predictions for Agentic AI in 2026. Whether you're an AI enthusiast, a business leader, or simply curious about the future of technology, this video is for you.
00:00:00 Introduction
00:01:28 What is Agentic AI?
00:12:20 Agentic AI vs Generative AI
00:33:33 Alexa+ Powered by Generative AI
00:44:47 Agentic AI Roadmap
00:51:15 Introduction to Artificial Intelligence
01:10:03 Introduction to Deep Learning
02:27:26 Artificial Neural Network
02:59:26 Transformers Explained Using Generative AI
03:07:05 Transformers Neural Networks Explained
03:14:29 What are Large Language Models?
03:35:28 What is Multimodal AI?
03:52:46 LLM vs SLM
03:58:14 What is LangChain?
04:15:29 Langchain Agents Explained
04:22:44 What is RAG?
04:45:32 LLMOps: The Future of AI Development
04:51:55 Prompt Engineering
05:06:09 Natural Language Processing (NLP) & Text Mining using NLTK
05:44:52 KNN Algorithm using Python
06:03:11 Alibaba’s Qwen 2.5-Max Just Beat GPT-4 & DeepSeek?
06:07:35 DeepSeek Training Cost: How China Built AI for Less
06:13:04 DeepSeek vs OpenAI: Who Wins the AI Race?
06:25:36 Deep Learning Interview Questions
✅Subscribe to our channel to get video u
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
Playlist
Uploads from edureka! · edureka! · 0 of 60
← Previous
Next →
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
ChatGPT Not Working - 4 Fixes | How To Fix ChatGPT Not Working | Why Is ChatGPT Not Working |Edureka
edureka!
Advanced Java script Tutorial | JavaScript Training | JavaScript Programming | Edureka Rewind
edureka!
Java script interview question and answers | Java script training | Edureka Rewind
edureka!
OpenAI API Tutorial using Python | How to use OpenAI GPT-3 API - Ada Babbage Curie Davinci | Edureka
edureka!
What is Unsupervised Learning ? | Unsupervised Learning Algorithms| Machine Learning | Edureka
edureka!
Top 10 Applications of Machine Learning in 2023 | Machine Learning Training | Edureka Rewind - 7
edureka!
Machine Learning Engineer Career Path in 2023 | Machine Learning Tutorial | Edureka Rewind - 6
edureka!
10 Must Have Machine Learning Engineer Skills That Will Get You Hired | Edureka Rewind - 7
edureka!
Data Structures in Python | Data Structures and Algorithms in Python | Edureka | Python Live - 5
edureka!
Python Lists | List in Python | Python Training | Edureka Rewind
edureka!
Predictive Analysis Using Python | Learn to Build Predictive Models | Python Training | Edureka
edureka!
Machine Learning Tutorial | Machine Learning Algorithm | Machine Learning Engineer Program | Edureka
edureka!
How to use Pandas in Python | Python Pandas Tutorial | Python Tutorial | Edureka Rewind
edureka!
Parameters in Tableau | Tableau Parameters Examples | Tableau Tutorial | Edureka Rewind
edureka!
Top 10 Reasons to Learn Tableau in 2023 | Tableau Certification | Tableau | Edureka Rewind
edureka!
Tableau Developer Roles & Responsibilities | Become A Tableau Developer | Tableau | Edureka Rewind
edureka!
Deep Learning With Python | Deep Learning Tutorial For Beginners | Edureka Rewind
edureka!
Realtime Object Detection | Object Detection with TensorFlow | Edureka | Deep Learning Rewind - 2
edureka!
Top 20 Tableau Tips and Tricks in 20 Minutes | Tableau Tutorial | Tableau Training | Edureka Rewind
edureka!
Climate Change Prediction using Time Series | Python Projects | Edureka | DS Rewind - 5
edureka!
ReactJS Installation Tutorial | ReactJS Installation On Windows | ReactJS Tutorial | Edureka Rewind
edureka!
Phases in Cybersecurity | Cybersecurity Training | Edureka | Cybersecurity Rewind - 2
edureka!
What Is React | ReactJS Tutorial for Beginners | ReactJS Training | Edureka Rewind
edureka!
Cybersecurity Frameworks Tutorial | Cybersecurity Training | Edureka | Cybersecurity Rewind- 2
edureka!
React vs Angular 4 | Angular 2 vs React | React & Angular | ReactJS Training | Edureka Rewind - 5
edureka!
ReactJS Components Life-Cycle Tutorial | React Tutorial for Beginners | Edureka Rewind
edureka!
Ethical Hacking using Kali Linux | Ethical Hacking Tutorial | Edureka | Cybersecurity Rewind - 3
edureka!
Types Of Artificial Intelligence | Artificial Intelligence Explained | What is AI? | Edureka
edureka!
Top 10 Applications Of Artificial Intelligence in 2023 | Artificial Intelligence| Edureka Rewind
edureka!
The Future of AI | How will Artificial Intelligence Change the World in 2023? | Edureka Rewind
edureka!
What is Artificial Intelligence | Artificial Intelligence Tutorial For Beginners | Edureka Rewind
edureka!
Google Cloud IAM | Identity & Access Management on GCP | Edureka | GCP Rewind - 5
edureka!
Google Cloud AI Platform Tutorial | Google Cloud AI Platform | GCP Training | Edureka Rewind
edureka!
Projects in Google Cloud Platform | GCP Project Structure | GCP Training | Edureka Rewind
edureka!
How to Become a Data Scientist | Data Scientist Skills | Data Science Training | Edureka Rewind - 3
edureka!
Agglomerative and Divisive Hierarchical Clustering Explained | Data Science Training | Edureka Live
edureka!
Climate Change Prediction using Time Series | Python Projects | Edureka | DS Rewind - 5
edureka!
Data Science Project - Covid-19 Data Analysis | Python Training | Edureka | DS Rewind - 6
edureka!
What is Honeycode? | Introduction to Honeycode | Edureka
edureka!
Difference between Amazon AWS and Google Cloud | GCP Training Google Cloud | Edureka Live
edureka!
DevOps Lifecycle | Introduction To DevOps | DevOps Tools | What is DevOps? | Edureka Rewind
edureka!
Introduction to DevOps | DevOps Tutorial for Beginners | DevOps Tools | DevOps | Edureka Rewind
edureka!
How to Create Login System using Python | Python Programming Tutorial | Edureka Rewind
edureka!
Python Developer | How to become Python Developer | Python Tutorial | Edureka Rewind
edureka!
How to become a Data Engineer | Complete Roadmap to become a Data Engineer| Data Engineer | Edureka
edureka!
Azure Data Engineer Certification [DP 203] | How to Become Azure Data Engineer [2023] | Edureka
edureka!
Data Analyst vs Data Engineer vs Data Scientist | Data Analytics Masters Program | Edureka Rewind
edureka!
DevOps Engineer day-to-day Activities | DevOps Engineer Responsibilities | Edureka Rewind
edureka!
How to Become a DevOps Engineer? | DevOps Engineer Roadmap | Edureka | DevOps Rewind
edureka!
How to Become a Data Engineer? | Data Engineering Training | Edureka
edureka!
How To Become A Big Data Engineer? | Big Data Engineer Roadmap | Edureka Rewind
edureka!
Python Integration for Power BI and Predictive Analytics | Power BI Training | Edureka
edureka!
Power BI KPI Indicators Tutorial | Custom Visuals In Power BI | Power BI Training | Edureka Rewind
edureka!
Apache HBase Tutorial For Beginners | What is Apache HBase? | Big Data Training | Edureka Rewind
edureka!
Big Data Hadoop Tutorial For Beginners | Hadoop Training | Big Data Tutorial | Edureka Rewind
edureka!
Big Data Analytics | Big Data Analytics Use-Cases | Big Data Tutorial | Edureka Rewind
edureka!
What Is Power BI? | Introduction To Microsoft Power BI | Power BI Training | Edureka Rewind
edureka!
Triggers in Salesforce | Salesforce Apex Triggers | Salesforce Tutorial | Edureka Rewind
edureka!
How To Become A Salesforce Developer | Salesforce For Beginners| Salesforce Training Edureka Rewind
edureka!
Java ArrayList Tutorial | Java ArrayList Examples | Java Tutorial | Edureka Rewind
edureka!
More on: Agent Foundations
View skill →Related Reads
📰
📰
📰
📰
Green Is Not Correct — field notes from an AI orchestrator, end of a long shift
Medium · LLM
LLM Security Guide: Prompt Injection and API Exploitation
Medium · Cybersecurity
LLM Security Guide: Prompt Injection and API Exploitation
Medium · LLM
This Article Features a 1-Minute Looped Transformer Demo, and It’s Mind-Blowing!
Medium · LLM
Chapters (24)
Introduction
1:28
What is Agentic AI?
12:20
Agentic AI vs Generative AI
33:33
Alexa+ Powered by Generative AI
44:47
Agentic AI Roadmap
51:15
Introduction to Artificial Intelligence
1:10:03
Introduction to Deep Learning
2:27:26
Artificial Neural Network
2:59:26
Transformers Explained Using Generative AI
3:07:05
Transformers Neural Networks Explained
3:14:29
What are Large Language Models?
3:35:28
What is Multimodal AI?
3:52:46
LLM vs SLM
3:58:14
What is LangChain?
4:15:29
Langchain Agents Explained
4:22:44
What is RAG?
4:45:32
LLMOps: The Future of AI Development
4:51:55
Prompt Engineering
5:06:09
Natural Language Processing (NLP) & Text Mining using NLTK
5:44:52
KNN Algorithm using Python
6:03:11
Alibaba’s Qwen 2.5-Max Just Beat GPT-4 & DeepSeek?
6:07:35
DeepSeek Training Cost: How China Built AI for Less
6:13:04
DeepSeek vs OpenAI: Who Wins the AI Race?
6:25:36
Deep Learning Interview Questions
🎓
Tutor Explanation
DeepCamp AI