Blog

  • Shaken and Stirred: How telecom industry is dealing with Robocalls

    Shaken and Stirred: How telecom industry is dealing with Robocalls

    “While there is no silver bullet in the endless fight against scammers, STIR/SHAKEN will turbo-charge many of the tools we use in our fight against robocalls: from consumer apps and network-level blocking, to enforcement investigations and shutting down the gateways used by international robocall campaigns.” Jessica Rosenworcel, acting Chairwoman, FCC (1)

    Call spoofing is a menace to phone users across the globe. The intention behind these calls varies from simply maximizing the chances that the receiver will pick up the call to being part of a grander fraud scheme to steal identities, financial details, and more.

    Fraudulent robocalls will cost customers US $40 billion in 2022, up from US $31 billion in 2021, according to a study by Juniper Research (2).

    The alarm is officially raised. In March 2020, the US-based Federal Communications Commission (FCC) directed the telecommunications industry to actively and innovatively seek ways to combat robocalling.

    One of the prominent standards to start with is STIR/SHAKEN. Short for “Secure Telephone Identity Revisited” (STIR) and “Signature-based Handling of Asserted information using toKENs” (SHAKEN), the framework provides mechanisms that avert caller ID spoofing. By verifying caller IDs, the framework helps receivers identify that a robocall is, in fact, a robocall versus a legitimate call (3). The deadline for communication service providers to fully implement STIR/SHAKEN was June 2021.

    The secret world of caller spoofing
    Curious about the acronym STIR/SHAKEN? Indeed, it is inspired by secret agent 007, James Bond, who famously prefers his Martinis shaken and not stirred. Since the STIR framework was developed earlier, SHAKEN had to be the obvious choice. Jim McEachern, a senior technology consultant with the ATIS, wittily remarks, “We tortured the English language until we came up with an acronym” (7).

    How can standards help? 

    The STIR/SHAKEN framework uses digital certificates based on common public-key cryptography techniques to ensure that the calling number of a telephone call is secure.

    Those providers that fail to implement STIR/SHAKEN and register themselves in the robocall mitigation database won’t be able to provide domestic voice traffic services (4).

    The protocols themselves reinforce the telco network’s ability to prevent caller ID spoofing. STIR is the actual technology that fights illegal spoofing using digital certificates that cross-check the accuracy and validity of a calling number. SHAKEN guides telcos on how to properly implement STIR technology within their networks.

     

    The STIR/SHAKEN workflow 

    • The originating telephone service provider receives a SIP INVITE.
    • The call source and the calling number are checked by the originating telephone service provider to determine how to attest to the validity of the calling number.
      • Full Attestation (A) — The service provider has authenticated the calling party, and they are authorized to use the calling number.
      • Partial Attestation (B) — The service provider has authenticated the call origination but cannot verify that the call source is authorized to use the calling number.
      • Gateway Attestation (C) — The service provider has authenticated from where it received the call but cannot authenticate the call source.
    • The originating service provider creates a ‘SIP identity header.’ This notes details like the dialer’s number, the number being dialed, timestamp, attestation type, origination identifier, etc.
    • These two information buckets or digital certificates – the SIP Invite and the SIP Identity header – are sent to the destination service provider that passes it to the verification service.
    • The digital certificates are verified against the public certificate repository through a complex process, after which successful verification deems that the dialing number is not spoofed or illegal.
    • The call is approved for the destination service provider, and the call completes its journey, reaching the final party.

    Shaking up a stir across the globe 

    STIR is a globally accepted standard that can be implemented in any country. On the other hand, SHAKEN is specific for the United States (6).

    America has been enthusiastic about its adoption of STIR/SHAKEN. Verizon, an American telco and one of the largest in the world, implemented the FCC’s industry mandate by March 2019. It has fully upgraded its wireless network to STIR/SHAKEN (5). This move has allowed Verizon to protect more than 78 million customers from over 13 billion spoofed calls!

    By 2021, many other top US mobile carriers followed suit in adopting STIR/SHAKEN, such as AT&T, T-Mobile, and US Cellular (1).

    At the heels of the US, other countries are evaluating the STIR/SHAKEN standards and their effectiveness in combating illegal call spoofing according to their unique needs.

    Take the case of Ofcom, UK’s communications regulator, which plans to fully retire copper lines and adopt VoIP from the PSTN by January 2025 as a step towards implementing STIR/SHAKEN. But since the UK does not have a national telephone number database of assigned numbers, NICC, a UK-based tech forum, has suggested a three-phased approach for transition. Other examples are Canada and France.

    Teething troubles with STIR/SHAKEN

    While STIR/SHAKEN can be lauded for its advantages, it also has shortcomings.

    This is why even though the deadline for implementing STIR/SHAKEN is behind us, completely getting rid of spoofed calls remains an uphill task.

    Here’s why:

    • No compulsion for smaller providers – Small providers having less than 100,000 subscribers are exempt from the FCC mandate; they qualify for a two-year extension. Certain other extensions have also been made for non-IP portions of provider networks, making them the preferred route for scammers now (6).
    • Copper landline wires – Landline phone networks have also not been able to deploy STIR/SHAKEN as copper landlines are unable to support this technology (1).
    • Not applicable on SMS – STIR/SHAKEN applies to only phone calls and not SMS services. Scammers can still bombard millions of unwitting users with illegally spoofed messages over SMS (9).
    • A costly avenue – Implementing STIR/SHAKEN is a costly affair.

    More to be done! 

    Telecom may be far from achieving spoof-free, networks but implementing STIR/SHAKEN standards are a step in the right direction.

    Scammers are quick to adapt to obstacles and are constantly seeking clever ways to bypass stronger security protocols. Nations and legitimate providers must work together to create interoperable standards that weed out the menace of illegal and malicious robocalling.

    Regulators like the FCC must study the effectiveness of STIR/SHAKEN standards, addresses concerns periodically, and find the right tech partners to make implementation seamless and cost-friendly.

    References

    1. https://arstechnica.com/tech-policy/2021/07/us-hits-anti-robocall-milestone-but-annoying-calls-wont-stop-any-time-soon/
    2. https://www.juniperresearch.com/press/robocall-fraud-to-cost-consumers-$40bn
    3. https://www.consumeraffairs.com/news/major-phone-carriers-confirm-theyve-met-the-fccs-mandate-on-robocall-protection-070121.html
    4. https://www.natlawreview.com/article/fcc-reminders-re-stirshaken-and-robocall-mitigation-database
    5. https://www.verizon.com/about/news/over-78-million-verizon-customers-protected-over-13-billion-unwanted-calls
    6. https://www.dwt.com/insights/2020/10/fcc-stir-shaken-robocall-mitigation-plan-deadline
    7. https://www.latimes.com/business/lazarus/la-fi-lazarus-robocalls-fcc-task-force-20170901-story.html

    Combating Robocalls with Multi-Tiered Detection and Prevention Approach

    Download the point of view

  • Top 7 AI and Data Analytics Trends to look forward in 2022

    Top 7 AI and Data Analytics Trends to look forward in 2022

    There’s no doubt that AI and analytics are already changing how businesses operate across different industries- whether through task automation, insight generation, or other use cases. In 2022, it will continue to transform the enterprises in the way they live and work, and leaders will shift their focus to adopting AI solutions that are more sustainable and scalable.

    What 2022 is going to look like for AI and Data Analytics?

    In 2021, many enterprises saw the adoption of AI in at least one function across major industries. According to McKinsey Global Survey 2021, 57 percent of respondents report AI adoption, up from 45 percent in 2020.

    In 2022, preparation is the key to AI and Analytics success. It may be tempting to push AI into legacy environments as quickly as possible; it would be wiser to adopt a more careful and thoughtful approach. AI is only as good as the data it can access, so shoring up both infrastructure and data management and preparation processes will play a substantial role in adopting future AI-driven initiatives.

    From a technology perspective, there are many discussions around low-code/no-code AI platforms and architectural approaches to analytics, like Data Fabric, composable data, and analytics, workforce augmentation, etc. Let’s look at some interesting statistics for this year.

    Data analysis

    Top AI and Data Analytics trends to keep you on the radar

    Many new developments and breakthroughs will continue to push the boundaries of what’s possible. Here are the key areas where those breakthroughs will occur in 2022:

    #1 Data Fabric will be a key foundation for the enterprise to establish a frictionless data journey

    As the data increases in volume and becomes increasingly complex, and digital business accelerates, data fabric creates an agile and data-centric environment that responds quickly to the fast pace of change. According to Gartner, the data fabric concept enables frictionless access to and sharing data in a distributed data environment. It consists of end-to-end data integration and management solution that unlocks the potential of the data and reduces data to insights journey from any environment- cloud, on-premises, or edge. It reduces the time for integration for design by 30%, deployment by 30%, and maintenance by 70% because the technology designs draw on the ability to use/reuse and combine different data integration styles.

    #2 Composable data and analytics fosters agility

    Enterprises have more than one standard tool for analytics and BI. So, introducing new technology or tool becomes a costly affair. Composable data and analytics use/reuse components from multiple data, analytics, and AI solutions to reduce costs, boost deployment speed, and encourage collaboration. It uses AI across business intelligence, data management, and predictive analytics, evolving the analytics capabilities of an organization and enabling leaders to connect data insights to business actions.

    #3 Improved decision intelligence for enterprise-wide decision support

    Decision intelligence uses emerging technologies such as AI/ML to process large amounts of data to quickly extract meaningful insights for enterprises needed to drive actions for the business. In 2022, decision intelligence has the potential to make assessments not only better, but also faster, given that machine-generated decisions can be processed at speeds that humans cannot achieve.

    #4 Rise of Low-code and No-code technologies

    According to Gartner, by 2023, over 50% of medium to large enterprises will have adopted low-code or no-code as one of their strategic application platforms. Also, it predicts that low-code platforms will be responsible for more than 65% of application development activity by 2024. The low-code and no-code technologies enable businesses to keep up with the rapidly changing technology landscape through innovation and by empowering business users and technical programmers to build applications with little to no coding.

    No-code/Low-code vs. Code-Heavy: What’s the difference?

    Data analysis

    #5 Ethical AI and Ethical Data Analytics become tangible

    As enterprises power AI advancements, the lack of governmental oversights has pushed the debate over the ethics of responsible AI to the fore. In 2022, we will see how ethical AI and ethical data analytics will continue to play a significant part in the simulation of innovation and economic growth, since more organizations will realize the need for responsible tech. Fairness of algorithms and data transparency are issues that will need to be addressed in the coming years as AI adoption is more widespread than ever. It will hopefully work its way towards policymakers as well.

    #6 AI will evolve more rapidly, expanding and impacting every business process

    While in 2021, most of the enterprises were still in the proof-of-concept phase of AI. 2022 will see a shift towards AI-first approaches. AI applications will be at the forefront of enterprise strategies. As AI/ML models become the norm, companies will expand AI to become every part of the department and impact every business process.

    #7 AI will become more widespread and accessible

    Previously in 2021, only big players such as Amazon, Google, Microsoft, etc., had the deep pockets to make AI/ML models a reality. In 2022, there will be more off-the-shelf technology to make AI/ML models more accessible, like readily available functionality to make applications talk, convert speech to text, and other industry-specific use cases. Also, modern workplaces are evolving with AI getting incorporated into their processes. Humans and machines will work alongside each other for quicker results. This creates a great combination of human innovation and machine intelligence.

    In 2022, AI and data analytics will not only be more prevalent but will also be more strategic. It will continue to be used to achieve productivity gains. In the coming years, AI will also be used to rethink and redesign products, services, business models, and overall strategy.

    This year, the challenges of integrating, cleaning, and processing data will continue. However, at the same time, there will be a flood of more generic AI and data-analytics platforms that will help replace manual tasks, freeing up data scientists for strategic tasks. As today’s enterprises strive to be data-driven and demand that the data be most efficient to provide a better experience, more enabling technologies will be available to ease the transition and adoption of AI across the organization.

    Do you agree with the points discussed in the article? If yes, which of the trends do you plan to adopt in your organization? Feel free to let us know your thoughts and comments in the section below if we’ve missed any important points.

    Derive Maximum Value From Your AI Investments

    Request Demo

  • Top 7 AI and Data Analytics Trends to look forward in 2022

    There’s no doubt that AI and analytics are already changing how businesses operate across different industries- whether through task automation, insight generation, or other use cases. In 2022, it will continue to transform the enterprises in the way they live and work, and leaders will shift their focus to adopting AI solutions that are more sustainable and scalable.

    What 2022 is going to look like for AI and Data Analytics?

    In 2021, many enterprises saw the adoption of AI in at least one function across major industries. According to McKinsey Global Survey 2021, 57 percent of respondents report AI adoption, up from 45 percent in 2020.

    In 2022, preparation is the key to AI and Analytics success. It may be tempting to push AI into legacy environments as quickly as possible; it would be wiser to adopt a more careful and thoughtful approach. AI is only as good as the data it can access, so shoring up both infrastructure and data management and preparation processes will play a substantial role in adopting future AI-driven initiatives.

    From a technology perspective, there are many discussions around low-code/no-code AI platforms and architectural approaches to analytics, like Data Fabric, composable data, and analytics, workforce augmentation, etc. Let’s look at some interesting statistics for this year.

    Data analysis

    Top AI and Data Analytics trends to keep you on the radar

    Many new developments and breakthroughs will continue to push the boundaries of what’s possible. Here are the key areas where those breakthroughs will occur in 2022:

    #1 Data Fabric will be a key foundation for the enterprise to establish a frictionless data journey

    As the data increases in volume and becomes increasingly complex, and digital business accelerates, data fabric creates an agile and data-centric environment that responds quickly to the fast pace of change. According to Gartner, the data fabric concept enables frictionless access to and sharing data in a distributed data environment. It consists of end-to-end data integration and management solution that unlocks the potential of the data and reduces data to insights journey from any environment- cloud, on-premises, or edge. It reduces the time for integration for design by 30%, deployment by 30%, and maintenance by 70% because the technology designs draw on the ability to use/reuse and combine different data integration styles.

    #2 Composable data and analytics fosters agility

    Enterprises have more than one standard tool for analytics and BI. So, introducing new technology or tool becomes a costly affair. Composable data and analytics use/reuse components from multiple data, analytics, and AI solutions to reduce costs, boost deployment speed, and encourage collaboration. It uses AI across business intelligence, data management, and predictive analytics, evolving the analytics capabilities of an organization and enabling leaders to connect data insights to business actions.

    #3 Improved decision intelligence for enterprise-wide decision support

    Decision intelligence uses emerging technologies such as AI/ML to process large amounts of data to quickly extract meaningful insights for enterprises needed to drive actions for the business. In 2022, decision intelligence has the potential to make assessments not only better, but also faster, given that machine-generated decisions can be processed at speeds that humans cannot achieve.

    #4 Rise of Low-code and No-code technologies

    According to Gartner, by 2023, over 50% of medium to large enterprises will have adopted low-code or no-code as one of their strategic application platforms. Also, it predicts that low-code platforms will be responsible for more than 65% of application development activity by 2024. The low-code and no-code technologies enable businesses to keep up with the rapidly changing technology landscape through innovation and by empowering business users and technical programmers to build applications with little to no coding.

    No-code/Low-code vs. Code-Heavy: What’s the difference?

    Data analysis

    #5 Ethical AI and Ethical Data Analytics become tangible

    As enterprises power AI advancements, the lack of governmental oversights has pushed the debate over the ethics of responsible AI to the fore. In 2022, we will see how ethical AI and ethical data analytics will continue to play a significant part in the simulation of innovation and economic growth, since more organizations will realize the need for responsible tech. Fairness of algorithms and data transparency are issues that will need to be addressed in the coming years as AI adoption is more widespread than ever. It will hopefully work its way towards policymakers as well.

    #6 AI will evolve more rapidly, expanding and impacting every business process

    While in 2021, most of the enterprises were still in the proof-of-concept phase of AI. 2022 will see a shift towards AI-first approaches. AI applications will be at the forefront of enterprise strategies. As AI/ML models become the norm, companies will expand AI to become every part of the department and impact every business process.

    #7 AI will become more widespread and accessible

    Previously in 2021, only big players such as Amazon, Google, Microsoft, etc., had the deep pockets to make AI/ML models a reality. In 2022, there will be more off-the-shelf technology to make AI/ML models more accessible, like readily available functionality to make applications talk, convert speech to text, and other industry-specific use cases. Also, modern workplaces are evolving with AI getting incorporated into their processes. Humans and machines will work alongside each other for quicker results. This creates a great combination of human innovation and machine intelligence.

    In 2022, AI and data analytics will not only be more prevalent but will also be more strategic. It will continue to be used to achieve productivity gains. In the coming years, AI will also be used to rethink and redesign products, services, business models, and overall strategy.

    This year, the challenges of integrating, cleaning, and processing data will continue. However, at the same time, there will be a flood of more generic AI and data-analytics platforms that will help replace manual tasks, freeing up data scientists for strategic tasks. As today’s enterprises strive to be data-driven and demand that the data be most efficient to provide a better experience, more enabling technologies will be available to ease the transition and adoption of AI across the organization.

    Do you agree with the points discussed in the article? If yes, which of the trends do you plan to adopt in your organization? Feel free to let us know your thoughts and comments in the section below if we’ve missed any important points.

    Derive Maximum Value From Your AI Investments

    Request Demo

  • An Introduction to efficient Hypermeter optimization for XGBoost model using Optuna

    An Introduction to efficient Hypermeter optimization for XGBoost model using Optuna

    Efficient Hyperparam

    Introduction :

    Hyperparameter optimization is the science of tuning or choosing the best set of hyperparameters for a learning algorithm. A set of optimal hyperparameter has a big impact on the performance of any machine learning algorithm. It is one of the most time-consuming yet a crucial step in machine learning training pipeline.

    A Machine learning model has two types of tunable parameter :

    • Model parameters
    • Model hyperparameters

    Efficient Hyperparam

    Model parameters vs Model hyperparameters (source)

    Model parameters are learned during the training phase of a model or classifier. For example :

    • coefficients in logistic regression or liner regression
    • weights in an artificial neural network

    Model Hyperparameters are set by user before the model training phase. For example :

    • ‘c’ (regularization strength), ‘penalty’ and ‘solver’ in logistic regression
    • ‘learning rate’, ‘batch size’, ‘number of hidden layers’ etc. in an artificial neural network

    The choice of Machine learning model depends on the dataset, the task in hand i.e. prediction or classification. Each model has its own unique set of hyperparameter and the task of finding the best combination of these parameter is known as hyperparameter optimization.

    For solving hyperparameter optimization problem there are various methods are available. For example :

    • Grid Search
    • Random Search
    • Optuna
    • HyperOpt

    In this post, we will focus on Optuna library which has one of the most accurate and successful hyperparameter optimization strategy.

    Optuna :

    Optuna is an open source hyperparameter optimization (HPO) framework to automate search space of hyperparameter. For finding an optimal set of hyperparameters, Optuna uses Bayesian method. It supports various types of samplers listed below :

    • GridSampler(using grid search)
    • RandomSampler(using random sampling)
    • TPESampler (using Tree-structured Parzen Estimator algorithm)
    • CmaEsSampler ( using CMA-ES algorithm)

    Use of Optuna for hyperparameter optimization is explained using Credit Card Fraud Detection dataset on Kaggle. The problem statement is to classify a credit card transaction fraudulent or genuine(binary classification). This data contains only numerical input variables which are PCA transformation of original features. Due to confidentially issues, the original features and more background information about the data are not available.

    In this case, we have used only a subset of the dataset to speed up the training time and to ensure the two different classes reach a perfectly balance. Here, the sampling method used is TPESampler . A subset of a dataset is shown in the figure below :

    Efficient Hyperparam

    A subset of Credit Card Fraud Detection dataset

    Importing required packages :

    import optuna
    from optuna import Trial, visualization
    from optuna.samplers import TPESampler
    from xgboost import XGBClassifier

    Following are the main steps involved in HPO using Optuna for XGBoost model:

    1. Define Objective Function :
      The first important step is to define an objective function. The objective should be to return a real value which has to minimize or maximize. In our case, we will be training XGBoost model and using the cross-validation score for evaluation. We will be returning this cross-validation score from an objective function which has to be maximized.
    2. Define Hyperparameter Search Space :
      Optuna supports five kind of hyperparameters distribution, which are given as follows :
    • Integer parameter: A uniform distribution on integers.
      n_estimators = trial.suggest_int(‘n_estimators’,100,500)
    • Categorical parameter: A categorical distribution.
      criterion = trial.suggest_categorical(‘criterion’ ,[‘gini’, ‘entropy’])
    • Uniform parameter: A uniform distribution in linear domain.
      subsample = trial.suggest_uniform(‘subsample’ ,0.2,0.8)
    • Discrete-uniform parameter: A discretized uniform distribution in linear domain.
      max_features = trial.suggest_discrete_uniform(‘max_features’, 0.05,1,0.05)
    • Loguniform parameter: A uniform distribution in log domain.
      learning_rate = trial.sugget_loguniform(‘learning_rate’ : 1e-6, 1e-3)

    The below figure shows the objective function and hyperparameter for our example.

    Objective Function

    1. Study Objective :
      We have to understand some important terminologies mentioned in their docs, which will make our work easier. These are given as follows :
    • Trial : A single call of the objective function
    • Study : An optimization session, which is a set of trails
    • Parameter : A variable whose value is to be optimized such as value of “n_estimators”

    The Study object is used to manage optimization process. Method create_study() returns a study object. A study object has useful properties for analyzing the optimization outcome. In method of create_study(), we have to define the direction of objective function i.e. “maximize” or “minimize” and sampler for example TPESampler(). After creating study, we can call Optimize().

    1. Best Trial and Result :
      Once the optimization process completed then we can obtain the best parameters value and the optimal value of the objective function.

    Best trial: score 0.9427118644067797,
    params {‘n_estimators’: 396, ‘max_depth’: 6, ‘reg_alpha’: 3, ‘reg_lambda’: 3, ‘min_child_weight’: 2, ‘gamma’: 0, ‘learning_rate’: 0.09041583301198859, ‘colsample_bytree’: 0.45999999999999996}

    1. Trail History :
      We can get the entire history of all the trial in the form data frame by just calling study.trails_dataframe().

    Efficient Hyperparam

    1. Visualizations :

    Efficient Hyperparam

    Photo by Isaac Smith on Unsplash

    Visualizing the hyperparameter search space can be very useful. From the visualization, we can gain some useful information on the interaction between parameters and we can see where to search next. The optuna.visualization module includes a set of useful visualizations.

    i) plot_optimization_history(study): plots optimization history of all trials as well as the best score at each point.

    Efficient Hyperparam

    Optimization History Plot

    ii) plot_slice(study): plots the parameter relationship as slice also we can see which part of search space were explored more.

    Efficient Hyperparam

    Slice Plot

    iii) plot_parallel_coordinate(study) : plots the interactive visualization of the high-dimensional parameter relationship in study and scores.

    Efficient Hyperparam

    Parallel Coordinate Plot

    iv) plot_contor(study): plots parameter interactive chart from we can choose which hyperparameter space has to explore.

    Efficient Hyperparam

    Contour Plot

    Overall, Visualizations are Amazing in Optuna !!

    Want to see AI/ML models in action?

    Build your first AI model for Free with our 14-day free trial of HyperSense AI studio

    Get Free Trial

  • Evolution from Interconnect to Partner Management

    Evolution from Interconnect to Partner Management

    What is Interconnect Business?

    Interconnect is a process for telecom operators to handle calls for other operators, thus allowing people using different networks to communicate in domestic and international scenarios. Point of Interconnection connects the physical interface between two telecom operators to connect their customers. If operator A and operator B are not interconnected partners, their customers would not call each other. So, to allow ease of communication, operators get into interconnecting agreements with each other, thus allowing excellent business opportunities for them.

    What is Interconnect Billing System?

    It is a centralized, automated solution that supports multi-party agreements with multiple service providers to use the network and facilitate the traffic routing between multiple networks like circuit-switched networks (e.g., PSTN) or computer networks (e.g., Internet). The traffic flow is regulated through the rules defined by the system’s rules according to the interconnect agreement between the operators. The interconnect billing solution is near real-time processing to allow business optimization and help achieve increased operational efficiency and network profitability. As the business in most of the cases is bidirectional, billing systems perform two essential tasks-

    • Managing Deals or Bilateral Bulk Agreements
    • Partner Settlements –the incoming and outgoing invoices are taken through a net-off process where the net payable or receivables are identified

    Why is there a need for a Partner Management Solution?

    5G is beyond pure connectivity. CSPs need to work with other partners to develop compelling product bundles that provide ancillary services and processes to help enterprises achieve a significant strategic shift. CSPs should use their 5G assets to create value with their partners.

    While doing so, it is essential to acknowledge that partnerships play a critical role in the continued business success of service providers. Enhancing partner profitability is critical to retaining the right partnerships. Low partner profitability due to lack of settlement processes and systems may lead to a loss of invaluable partners, affecting service levels, revenues, and margins.

    Earlier, Voice revenues used to be the bread and butter for every Telco, but the emergence of Internet and OTT services have impacted revenues drastically. Gone are the days when operators entered into interconnecting partnerships only with other telecom operators to share and leverage each other’s network, primarily for Voice traffic and Messaging revenues. The proliferation of smart devices and high-speed internet has completely revolutionized the telecom ecosystem. A sudden rise in the plethora of partner-enabled services has led to a more significant divergence between partner management and interconnection systems. The number of interconnect partners has relatively increased. Still, the nature of partnerships has been relatively the same, which does not stand true for partner settlements where new partnerships are budding up quickly.

    Key differences between Interconnection and Partner Management Solution:

    Interconnect and Settlements

    • It mainly includes systems that record the volume and value of traffic. Primarily related to the rating for voice, data, and messaging
    • Traditional and widely deployed. Almost all mobile operators have some form of interconnection and settlement system in use
    • The revenue model is typically based on the volume of transactions
    • Settlements are mostly related to inter CSP connectivity

    Partner Management

    • It includes systems that enable third-party providers to offer their services to CSP customers and enable CSPs to charge for their service
    • Comparatively new and driven mainly by demand for partner-enabled digital economy services
    • The revenue model is mostly revenue sharing and, in some cases, is tied to product licensing
    • Support complex third-party settlements, subscription, billing, etc

    The wholesale telecom market is expanding aggressively, competition is fierce, margins have dwindled, billions of transactions/events take place which must be rated and charged, and multi-party agreements have become complex leading to a decrease in the quality of service rendered.

    There is a need for a Partner Management system that overcomes all CSP challenges for superior partner relations. This is where Subex Partner Management fits in. It is a convergent platform that offers a 360⁰ view of the evolving telecom ecosystem across Mobility, Content, and Entertainment, 5G for Business, Enterprise, and Internet of Things by providing a nuanced profile of partner agreements based on data such as revenue and margins. It helps in swift partner onboarding, partner self-care, partner assurance, end-to-end revenue visibility, and accessible communication between Telco and its partners. It is designed to co-exist with your legacy systems and manage diverse revenue streams while helping you launch high-value, high-margin services in collaboration with partners.

    From Effective to Cutting-edge: Next Generation Partner Lifecycle Management

    Download the point of view

  • Land record-keeping over blockchain

    Land record-keeping over blockchain

    Being a second populous country and seventh-largest nation globally, Land Record keeping in India is a vast transactional activity where most processes are manual and non-digital. Digital access of records, simplification of the process to bring transparency for buyer-to-seller to-authority, and a tamper-proof ledger are extremely necessary to deal with the increasing numbers of litigation cases.

    The current set-up of land records transactions involves multiple departments like land, revenue, forest, banks, and others as key stakeholders. These numerous stakeholders are involved in the transactions, so keeping track of deadlines and SLAs is difficult and complex, leading to delay in assigning ownership or transfer of ownership, etc. Any errors in the ledger can affect ownership rights. These manual ledgers and paperwork can be easily forged, affecting the ownership and leading to boundary or land litigation disputes.

    In the absence of a will, the partition of property among successors of the deceased owner is again a complex process. Even the validation and verification of forged documents during the process is not an easy task, and it leads to delay due to complexity even it’s not accurate.

    A blockchain-based next-generation end-to-end solution is required to overcome all these issues, ensuring privacy, security, and a tamper-proof platform.

    How can a decentralized blockchain-based solution help overcome the above challenges?

    1. Multiple vital stakeholders can be onboarded on the chain to share the data through smart contracts over the private channel as and when required quickly.
    2. SLA of each stakeholder or department can be tracked easily and made available in a real-time dashboard.
    3. A golden record or ownership database can be maintained, accessible to authorized users over the chain to validate the transactions and records.
    4. This mechanism can identify fake documents by verifying them against the golden database.
    5. Dynamic parameters like ownership, mortgage information, litigation status, property tax details, etc., can be easily derived from a tamper-proof master source.
    6. Any land dispute associated with boundaries can be immediately flagged in real-time.
    7. The sanction of loans against land or property can be facilitated easily and quickly.
    8. The allocation process of new lands or property will be more transparent.
    9. Transfer of ownership can be done over an immutable platform that can quickly be backtracked and remain tamper-proof.
    10. Once land allocation or transfer of ownership is confirmed, it can be approved on distributed ledger where each transaction is signed off with a digital signature, timestamp, and digital key of the user.

    A next-generation E2E solution over blockchain would be digital, decentralized, transparent, and immutable. It could help simplify the land record-keeping process where multiple departments are involved by bringing transparency across the chain by utilizing blockchain’s key features. A tamper-proof distributed ledger among the stakeholders will increase efficiency by improving response time, real-time analysis, and data availability.

    Streamlining partnership through decentralized technology

    Schedule Demo

  • Lessons from a Spy: How 007 is Helping Telecom Industry Fight Robocalling Fraud

    Lessons from a Spy: How 007 is Helping Telecom Industry Fight Robocalling Fraud

    In the iconic spy world brought to us through the imagination of Ian Fleming, I’ve found the array of gadgets almost as thrilling as the action sequences. Remember the invisible car in ‘Die Another Day’ and how James Bond was able to control it remotely. A little out of whack, right? But that was 20 years ago. One look at today’s autonomous cars, and somehow that level of innovation doesn’t seem so implausible anymore.

    Here’s another less known fact: In the time it has taken you to read this, nearly 55,000 robocalls have made its ways to phone numbers across the globe. Robocalls are automatically dialed telemarketing calls that play pre-recorded ads and promote products. Sounds harmless, right? Not always. In many cases, robocalling acts as a gateway to scam people out of millions of dollars, mostly through CLI spoofing.

    According to research, robocall fraud may cost consumers US $40 billion globally by 2022.

    Shaken, not stirred: A cocktail of safeguards

    In 2019, the Federal Communications Commission was called in to help. They came up with the Secure Telephone Identity Revisited (STIR) and Signature-based Handling of Asserted information using toKENs, also known as the STIR/SHAKEN framework.

    Admittedly, the FCC had to do a lot of head-scratching to come up with a name that would render the acronym STIR/SHAKEN. If the acronym sounds familiar, it is because it was inspired by secret agent 007 James Bond himself who famously prefers his Martinis shaken, not stirred.

    STIR/SHAKEN is a set of rules, protocols, and procedures designed to enhance call integrity through authenticating caller ID information by assigning each call with an encrypted digital fingerprint, enabling receivers to tell an illegally spoofed call from a legitimate one. In March 2020, the framework became a mandate for all CSPs to follow to curb the issue is CLI spoofing and, in effect, robocalls.

    While US regulators are busy enforcing STIR/SHAKEN for the common good, there are some loopholes:

    • STIR/SHAKEN presently only works with IP-based telephone networks. Service providers will not be able to properly authenticate calls originating from non-IP systems such as copper landline wires.
    • The authentication process does not indicate whether a call is legal/illegal or wanted/unwanted. It just digitally attests to whether the caller can use the particular number.
    • STIR/SHAKEN applies to only phone calls but not to text messaging. Scammers can still originate illegal messages via spam SMS
    • Finally, the framework is expensive to implement, making smaller carriers shy away from adopting these standards.

    STIR/SHAKEN is indeed the right step forward in reducing robocalls for good, but there is still more work to be done. The issue of crime through robocalling is heavily disguised. Cybercriminals are moving from people to faceless, nameless robots. The i3 Forum report called ‘Caller ID Spoofing’ believes that relying on industry standards alone may not be enough to fight the villain in robocalls. The need of the hour is advanced high-tech real-time capabilities relying on analysis, investigation, probabilities, and a deep dive into patterns of fraud.

    Why regulators and CSPs must work together

    James Bond aficionados already know Agent Q – Bond’s go-to person at the research and development division of the British Secret Service. Q is known for equipping Bond with state-of-the-art tech to fend off villains. Bond is often incredulous about Q’s inventions until they save his life many times in the field. The tempering force, or regulator, in this seemingly tenuous relationship between the technology-savvy Quartermaster and the expert spy, is the Agent M. I find that this trifecta mirrors what we see between regulators, CSPs, and technology. To effectively reduce robocalling fraud, both entities – telcos and regulators – must work together, incorporating the latest technologies.

    And, STIR/SHAKEN is one step in the right direction.

    On their part, US regulators have released a roadmap of anti-robocall principles that serve as a guide for CSPs. They recommend:

    • Enable call blocking and call labeling services to all customers at no charge
    • Implement STIR/SHAKEN call authentication
    • Monitor network traffic, especially high-volume calls, to gauge patterns similar to robocalls
    • Investigate suspicious calls, identify the source, and institute ways to terminate these calls.
    • Confirm the identity of commercial customers to whitelist these
    • Strictly enforce ‘traceback’ so that illegal robocalls can be checked during the transport of voice calls

    Fraud Management Solutions are Telecom’s Agent Q

    Although STIR/SHAKEN is an excellent starting point for the CSPs to address the robocall menace, it is not enough to ensure an end to robocalls. Using analytics in addition to STIR/SHAKEN will provide insights for quick action.

    In fact, in 2019, the FCC allowed telcos to block calls based on reasonable analytics designed to identify unwanted calls without explicit consumer action as long as the consumer is given the option to opt-out of the blocking service. Additionally, in July 2020, the FCC further strengthened analytics-based blocking by adding safe harbors for service providers from liability under the Communications Act and the Commission’s rules for erroneous call blocking. It is now critical to have an advanced analytical solution that provides a multi-tier defense mechanism to combat this. Solutions such as the Subex Fraud Management system makes use of real-time signaling level analysis, paired with advanced machine learning techniques and hybrid rule engine, thus providing new opportunities to drive a prevention-based approach.

    CSPs should urgently take up the issue of robocalling as a priority. Its long-standing impact on customers can be rather severe. Robocall fraud can cause operators to completely lose customer trust and credibility, resulting in substantial revenue loss.

    Time is of the essence. In this respect, reel life mimics real life. As Agent Q famously tells James Bond, “I can do more damage on my laptop, sitting in my pajamas, before my first cup of Earl Grey than you can do a year in the field.”

    Combating Robocalls with Multi-Tiered Detection and Prevention Approach

    Download the point of view

  • Clustering Algorithms Part 1

    Clustering Algorithms Part 1

    To understand clustering, we need to have a basic knowledge of Machine Learning. Machine learning is a subset of Artificial Intelligence that allows a machine to automatically learn from past data without programming explicitly. Classical machine learning is often categorized by how an algorithm learns to become more accurate in its predictions. There are four basic approaches: supervised learning, unsupervised learning, semi-supervised learning, and reinforcement learning. The type of algorithm that data scientists choose to use depends on what type of data they want to predict. Supervised learning is a machine learning approach that’s defined by its use of labelled datasets. These datasets are designed to train or “supervise” algorithms into classifying data or predicting outcomes accurately. Unsupervised learning on the other hand deals with unlabelled datasets. Clustering is an application of unsupervised learning. Semi-supervised learning is a branch of machine learning that attempts to solve problems that require or include both labelled and unlabelled data. Semi-supervised learning employs concepts of mathematics such as characteristics of both clustering and classification methods. Reinforcement learning is a kind of Machine Learning where in the system that is to be trained to do a particular job, learns on its own based on its previous experiences and outcomes while doing a similar kind of a job.

    What is Clustering and How it Works?

    Clustering is the task of dividing the population or data points into several groups such that data points in the same groups are similar to other data points in that group and dissimilar to the data points in other groups. It is basically an assembly of objects based on similarity and dissimilarity between them.

    clustering blog image

    The Importance of Clustering

    Clustering helps in understanding the natural grouping in a dataset. Their motivation is to check out to parcel the information into some gathering of legitimate groupings. Grouping quality relies upon the strategies and the identification of hidden patterns. The biggest advantage of clustering over-classification is it can adapt to the changes made and helps single out useful features that differentiate different groups.

    The Usage of Clustering Algorithms in Real World

    It is widely used in many applications such as image processing, data analysis, and pattern recognition.

    It can be used in the field of biology, by deriving animal and plant taxonomies, identifying genes with the same capabilities.

    It also helps in information discovery by classifying documents on the web.

    It helps marketers to find the distinct groups in their customer base and they can characterize their customer groups by using purchasing patterns.

    Different Types of Clustering Methods

    Connectivity-based Clustering (Hierarchical clustering)

    Hierarchical Clustering is a method of unsupervised machine learning clustering where it begins with a pre-defined top to bottom hierarchy of clusters. It then proceeds to perform a decomposition of the data objects based on this hierarchy, hence obtaining the clusters

    Centroids-based Clustering (Partitioning methods) 

    Centroid based clustering is considered as one of the simplest clustering algorithms, yet the most effective way of creating clusters and assigning data points to it. The intuition behind centroid-based clustering is that a cluster is characterized and represented by a central vector and data points that are in close proximity to these vectors are assigned to the respective clusters.

    Distribution-based Clustering

    Distribution-based clustering creates, and groups data points based on their likely hood of belonging to the same probability distribution in the data

    Density-based Clustering (Model-based methods)

    Density-based clustering methods take density into consideration instead of distances. Clusters are considered as the densest region in a data space, which is separated by regions of lower object density, and it is defined as a maximal set of connected points.

    Fuzzy Clustering

    The general idea about clustering revolves around assigning data points to mutually exclusive clusters, meaning, a data point always resides uniquely inside a cluster, and it cannot belong to more than one cluster. Fuzzy clustering methods change this paradigm by assigning a data-point to multiple clusters with a quantified degree of belongingness metric.

    Constraint-based (Supervised Clustering)

    The clustering process, in general, is based on the approach that the data can be divided into an optimal number of “unknown” groups. The underlying stages of all the clustering algorithms to find those hidden patterns and similarities, without any intervention or predefined conditions

    If you are working with ML algorithms, chances are you will be widely using Clustering. Clustering is an incredibly useful unsupervised machine learning method that has a wide variety of applications.

    Get ahead with MLOps. Get better, faster results from your data.

    Try HyperSense AI Studio for Free

  • Introduction to clustering in data science

    To understand clustering, we need to have a basic knowledge of Machine Learning. Machine learning is a subset of Artificial Intelligence that allows a machine to automatically learn from past data without programming explicitly. Classical machine learning is often categorized by how an algorithm learns to become more accurate in its predictions. There are four basic approaches: supervised learning, unsupervised learning, semi-supervised learning, and reinforcement learning. The type of algorithm that data scientists choose to use depends on what type of data they want to predict. Supervised learning is a machine learning approach that’s defined by its use of labelled datasets. These datasets are designed to train or “supervise” algorithms into classifying data or predicting outcomes accurately. Unsupervised learning on the other hand deals with unlabelled datasets. Clustering is an application of unsupervised learning. Semi-supervised learning is a branch of machine learning that attempts to solve problems that require or include both labelled and unlabelled data. Semi-supervised learning employs concepts of mathematics such as characteristics of both clustering and classification methods. Reinforcement learning is a kind of Machine Learning where in the system that is to be trained to do a particular job, learns on its own based on its previous experiences and outcomes while doing a similar kind of a job.

    What is Clustering and How it Works?

    Clustering is the task of dividing the population or data points into several groups such that data points in the same groups are similar to other data points in that group and dissimilar to the data points in other groups. It is basically an assembly of objects based on similarity and dissimilarity between them.

    clustering blog image

    The Importance of Clustering

    Clustering helps in understanding the natural grouping in a dataset. Their motivation is to check out to parcel the information into some gathering of legitimate groupings. Grouping quality relies upon the strategies and the identification of hidden patterns. The biggest advantage of clustering over-classification is it can adapt to the changes made and helps single out useful features that differentiate different groups.

    The Usage of Clustering Algorithms in Real World

    It is widely used in many applications such as image processing, data analysis, and pattern recognition.

    It can be used in the field of biology, by deriving animal and plant taxonomies, identifying genes with the same capabilities.

    It also helps in information discovery by classifying documents on the web.

    It helps marketers to find the distinct groups in their customer base and they can characterize their customer groups by using purchasing patterns.

    Different Types of Clustering Methods

    Connectivity-based Clustering (Hierarchical clustering)

    Hierarchical Clustering is a method of unsupervised machine learning clustering where it begins with a pre-defined top to bottom hierarchy of clusters. It then proceeds to perform a decomposition of the data objects based on this hierarchy, hence obtaining the clusters

    Centroids-based Clustering (Partitioning methods) 

    Centroid based clustering is considered as one of the simplest clustering algorithms, yet the most effective way of creating clusters and assigning data points to it. The intuition behind centroid-based clustering is that a cluster is characterized and represented by a central vector and data points that are in close proximity to these vectors are assigned to the respective clusters.

    Distribution-based Clustering

    Distribution-based clustering creates, and groups data points based on their likely hood of belonging to the same probability distribution in the data

    Density-based Clustering (Model-based methods)

    Density-based clustering methods take density into consideration instead of distances. Clusters are considered as the densest region in a data space, which is separated by regions of lower object density, and it is defined as a maximal set of connected points.

    Fuzzy Clustering

    The general idea about clustering revolves around assigning data points to mutually exclusive clusters, meaning, a data point always resides uniquely inside a cluster, and it cannot belong to more than one cluster. Fuzzy clustering methods change this paradigm by assigning a data-point to multiple clusters with a quantified degree of belongingness metric.

    Constraint-based (Supervised Clustering)

    The clustering process, in general, is based on the approach that the data can be divided into an optimal number of “unknown” groups. The underlying stages of all the clustering algorithms to find those hidden patterns and similarities, without any intervention or predefined conditions

    If you are working with ML algorithms, chances are you will be widely using Clustering. Clustering is an incredibly useful unsupervised machine learning method that has a wide variety of applications.

    Get ahead with MLOps. Get better, faster results from your data.

    Try HyperSense AI Studio for Free

  • Overview of Tree-based Algorithms in under 10 Minutes

    Overview of Tree-based Algorithms in under 10 Minutes

    Machine learning is not all about AI, but it is a big part of it.

    Machine Learning has emerged as a new way of communicating your wishes to a computer. It’s exciting because it allows you to automate the ineffable. Machine learning is powering most of the recent advancements in AI, including Computer Vision, NLP (Natural Language Processing), Predictive Analytics, Chatbots, and a wide range of applications. To move up the data value chain from the information level to the knowledge level, we need to apply machine learning that will enable systems to identify patterns in data and learn from those patterns to apply to new, never-seen data.

    Difference between Supervised Learning, Unsupervised Learning, and Semi-Supervised Learning

    Machine Learning use cases primarily fall into 3 categories: Supervised Learning, Unsupervised Learning, and Semi-Supervised Learning.

    1) Supervised Learning

    In the Supervised Learning use case, we have a known set of input data and the corresponding set of responses for it. The aim is to build a model which aims to learn the patterns and relationships between different input features (called as training phase) to generate a reasonable prediction as a response to the new input data (called as testing phase). Most of the real-world problems at scale, addressed by different organizations, fall in this category.

    2) Unsupervised Learning

    Unlike the Supervised Learning use case, Unsupervised learning involves finding inferences and hidden patterns from input data without references to any labeled responses or outcomes. This ability to discover patterns in information makes it an ideal solution for Cross-Selling strategies, Customer Segmentation, Image and Pattern recognition.

    3) Semi-Supervised Learning

    Semi-Supervised Learning, on the other hand, offers solutions that use the best of both worlds: Supervised and Unsupervised Learning. This approach is mostly adopted when there is an absence of good quantity or quality input data with labeled outcomes to train a supervised learning model. The smaller labeled data set is used to identify hidden patterns and the larger unlabelled dataset is used to extract features to improve upon the lack of quantity of labeled dataset.

    What are Tree-based Algorithms?

    Tree-based algorithms are one of the best and most used supervised learning methods, this is mainly due to the reason that these predictive models offer high accuracy, stability, and ease of interpretation as they map non-linear relationships quite well. This series is a journey designed to help readers not only to get a peek under the hood of popularly used state of the art tree-based algorithms but also to help readers to reach a level of proficiency so that one can make a better choice at adopting them to solve the problem statement at hand. This blog specifically provides an overview of the following Tree-Based Algorithms, which are popularly used in the ML community:

    1. Decision Trees
    2. Random Forests
    3. Gradient Boosting Machines
    4. XGBoost
    5. LightGBM
    6. CatBoost

    1) Understanding Decision Trees

    Decision Trees are one of the simplest and intuitive ways of predictive modeling. As the name suggests, we construct a tree-like model of decisions. It’s a flowchart-like structure that performs several tests on different features of the data point. Based on the tests, the data point will either be assigned to a class or a continuous value outcome. That is to say, Decision Trees can be used both for Classification and Regression use cases.

    But how are Decision Trees constructed, in the first place? There are various algorithmic approaches to construct a decision tree-like ID3, C4.5, CART, etc. The main parameter which separates these different algorithms is the Splitting Criterion which is used while selecting the decision nodes of the Tree. The splitting criterion helps in deciding which feature to use and the threshold value to perform the decision-making.

    We want the Decision Tree Model to be small and generalized to the training set, hence the main intent while performing the splits using Decision Nodes is to obtain the purest child nodes. It’s the different ways of measuring this purity which has let us come up with different splitting criterion.

    Information Gain is one such criterion. For each node of the tree, we measure how much information the feature gives us about the class. The split with the highest information gain will be taken first and the process continues until all the child nodes are pure, or the information gain is 0. Information Gain is calculated by measuring the difference between the entropy of the original dataset before and after the split. ID3 uses information gained for constructing the decision trees.

    Decision Trees essentially combine complex rules to give the correct predictions, because of this design, it is prone to Overfitting. In other words, A small change in the training set can cause a large change in the structure of the decision tree causing instability. In Addition, using a Decision Tree is relatively expensive, compared to other tree-based models, when training a large dataset.

    Hence, a Decision Tree becomes a good choice when we have a small to medium-sized dataset and model interpretability is required. In other cases, it proves to be inefficient and hence other tree-based algorithms are preferred which we will explore in the next section.

    2) Understanding Random Forest

    Random Forest is essentially inspired to solve the overfitting problems in Decision Trees. As the name suggests, this model consists of many individual Decision Trees whose predictions are aggregated (Averaging in case of Regression and Max Voting in case of Classification).

    The idea behind aggregating the prediction is quite simple, it is based on the Wisdom of Crowds. It says that A large crowd with uncorrelated opinions combined provides more wisdom than an individual opinion. In Data Science terminology, A large number of relatively uncorrelated models operation as a group will outperform any individual constituents. Uncorrelated models protect from propagating the errors of the base model. While some trees might be wrong, other trees might be correct, and overall as a group, they learn the correct patterns from the data instead of overfitting like a decision tree.

    But how does the Random Forest Ensure to train individual decision trees which are uncorrelated? A simple answer to this would be, to supply different samples of data to each Decision Tree and grow each Decision Tree using a specific subset of features. This technique to randomly sample with replacement to build individual trees is also known as Bootstrap Aggregation or Bagging. Bagging in combination with Feature Randomness ensures that each tree is trained on different data and a different subset of features, ensuring that results are uncorrelated.

    Although Random Forest was successful in mitigating the overfitting problem (or High Variance issue) in Decision Trees it is at a cost of increased compute and resources. Since multiple Decision Trees need to be built and maintained for predictions, it tends to take longer time, higher compute, larger resources, and lacks interpretability because of the ensemble nature. Despite these drawbacks, the model’s efficiency in giving correct predictions and ability to be used both for Regression and Classification problems has led to it being a popular algorithm that is used across multiple domains.

    3) Understanding Gradient Boosting Machines

    To understand GBM, we first need to touch upon the concept of Boosting. Unlike Bagging, where predictions of base models are aggregated to come up with a final prediction, Boosting focuses on building a strong learner by iteratively or sequentially learning patterns from the data using a consecutive chain of weak learners. Each tree, in this chain, focuses on correcting the net error generated from the previous tree. The first tree is built on the features with the highest predictive power which then passes on the predictions and error (difference between the actual and prediction) as an input to subsequent tree along with a weight parameter. The weights are used to control the second tree to utilize only those features which will fine-tune the combined predictions of the first and second tree to be close to the target as much as possible. The final prediction of the GBM is the weighted sum of all the individual predictions.

    Now that we know how Boosting works, the question arises on how we come up with the optimal weights to be assigned to the predictions of the individual weak learner. This is where the Gradient Descent Algorithm comes in place. In the context of GBM, we want to build an additive model, wherein Trees are added one at a time while existing trees in the model are not changed. Gradient Descent aims to add trees that reduce the loss function (and follow the gradient). The weights assigned to each tree are updated to minimize the error. We can limit the number of trees to be added by fixing a number or acceptable threshold of the loss function to stop the training.

    Because of this design, Gradient Boosting is a Greedy Algorithm and hence will keep on improving to minimize all errors leading to overfitting. Additionally, because trees must be trained sequentially it is computationally expensive and memory exhaustive for a larger dataset. That said, GBM is great at handling complex, non-linear relationships and are more powerful and accurate as compared to Random Forest or Decision Trees. In addition, Data Pre-processing steps to be done before using GBM are minimal, as it Handles the Missing Data and works great with Categorical Data as well.

    4) Understanding XGBoost

    XGBoost or Extreme Gradient Boosting is a modification of the GBM Architecture designed for Speed and Performance. It is an effort that is directed to push the limits of computations resources for boosted tree algorithms. Because of the Execution Speed and Model Performance, this is one of the go-to models in hackathons and has been recently dominating the applied machine learning community and Kaggle competitions which requires building ML models on structured or tabular data.

    But what makes XGBoost so special? It’s the System Optimization and Algorithmic Enhancements on top of the existing GBM framework, which makes it stands out from the rest ML algorithm. Firstly, XGBoost approaches the Sequential Tree Building using Parallelized implementation. In addition, it addresses the overfitting problem in GBM by using the “depth-first” approach to pruning trees backward which adds to the computational performance. Secondly, The library is designed with cache awareness and out-of-core computing which optimizes the disk space while handling large datasets. Lastly, The Algorithmic Enhancements such as built-in Cross-Validation, Sparsity Awareness, and Regularization make it robust and enable it to deliver good model performance.

    5) Understanding LightGBM

    Like XGBoost, this is an open-source library that provides more efficient and effective implementation of the Gradient Boosting Algorithm. This Tree-Based model can be used both for Classification and Regression. This extends the idea of GBM by adding in a type of Auto Feature Selection as well as focusing on data points with large gradients. This results in a dramatic speed-up of training (Hence, it is called a lighter version of GBM) and predictive performance.

    The high efficiency in terms of computing and high model performance can be attributed to two novel techniques GOSS or Gradient-Based One-Side Sampling and EFB or Exclusive Feature Bundling. GOSS is a modification of the Gradient Boosting method which gives more weightage to those examples which results in a larger gradient, which speeds up the learning process. EFB is an approach to bundle sparse (mostly zero) categorical features which have been one-hot encoded, thus acting as an Automatic Feature selection method.

    But how does it compare against XGBoost? LightGBM has many of the XGBoost’s advantages such as Sparse Optimization, Parallel Training, Regularization, Early Stopping, etc. But a major difference lies in the way the trees are constructed. Light GBM grows trees Leaf-wise, unlike other Tree Ensemble methods which grow trees level-wise row by row. The leaf which leads to the largest decrease in loss is selected and a Tree is grown from the output of that Leaf.

    Although it’s not fair to compare LightGBM and XGBoost, prior works which tried to benchmark the optimizations in these modified versions of GBM suggest XGBoost as a powerful algorithm that reduces the maximum training time ( when used on a GPU). It was also seen that LightGBM is not fast enough to converge to a good set of hyperparameters, as compared to XGBoost. That said, both these algorithms enjoy a fair bit of popularity in the ML community when it comes to Hackathons and working on use cases with a large dataset.

    6) Understanding CatBoost

    The name CatBoost comes from two words “Category” and “Boosting”. CatBoost is an open-source machine learning algorithm developed by Yandex. This algorithm is built on top of a Gradient Boosting Architecture, wherein consecutive trees are spawned to decrease the net error in prediction, the only difference being CatBoost uses oblivious decision trees to grow a balanced tree. That is to say that it uses the same features to make left and right splits for each level of the tree. This design helps in more efficient usage of CPU, in turn reducing the training time.

    One of the key highlights of this algorithm is the ease of dealing with categorical features in the dataset. Almost all the ML models which have been built require all the training data to be numeric. That means, if we have a categorical feature, we will have to employ Label Encoding or One Hot Encoding technique, else it would fail during the model building stage. CatBoost, on the other hand, doesn’t require explicit pre-processing to convert the categories, instead it internally uses various statistical measures to combine different categorical features and numerical features to achieve the same. It is also important to note that providing One Hot Encoded inputs to CatBoost models can decrease its efficiency in terms of model performance.

    As far as the performance is concerned, this model provides state-of-the-art results and stands in the same league as other Boosting Algorithms like XGBoost, LGBM, etc. The ease of use and the robustness towards overfitting leads to building a more generalized model which performs well for both regression and classification problems.

    Conclusion

    I hope by now, you have developed a fair bit of understanding about the evolution of Tree-Based Algorithms in the Machine Learning Landscape along with the key features which differentiate each of them. Stay tuned for more such “under 10 minutes” series of blogs, in the same space, where we will delve deeper into each of the algorithms with an ML use case to understand the intricacies to be taken care of while using them.

    Build machine learning models within minutes

    Claim your free trial now