We use cookies, including third-party cookies from Google to serve personalized ads through AdSense, to operate this site and understand how it is used. By continuing to browse, you accept this use. See our Privacy Policy and Terms of Use for details, including how to opt out of personalized advertising.
Accept
SmartData CollectiveSmartData Collective
  • Analytics
    AnalyticsShow More
    chatgpt image jul 21, 2026, 04 34 30 pm
    4 Core Benefits of Predictive Maintenance after Vibration Analysis
    10 Min Read
    How Does Data Mining Boost Customer Satisfaction in Logistics? Harnessing Analytics for Results -- AI-generated illustration
    How Does Data Mining Boost Customer Satisfaction in Logistics? Harnessing Analytics for Results
    11 Min Read
    chatgpt image jul 13, 2026, 04 23 45 pm
    How Data Analytics Helps Companies Improve User Engagement
    19 Min Read
    chatgpt image jul 13, 2026, 03 59 46 pm
    How Data Analytics Improves Multi-Location Search Strategies
    10 Min Read
    cybersecurity efforts
    How Behavioral Analytics and AI Are Redefining Cybersecurity for Boca Raton Businesses
    14 Min Read
  • Big Data
  • BI
  • Exclusive
  • IT
  • Marketing
  • Software
Search
© 2008-25 SmartData Collective. All Rights Reserved.
Reading: The Statistics of Everyday Talk
Share
Notification
Font ResizerAa
SmartData CollectiveSmartData Collective
Font ResizerAa
Search
  • About
  • Help
  • Privacy
Follow US
© 2008-23 SmartData Collective. All Rights Reserved.
SmartData Collective > Uncategorized > The Statistics of Everyday Talk
Uncategorized

The Statistics of Everyday Talk

ThemosKalafatis
ThemosKalafatis
5 Min Read
The Statistics of Everyday Talk
Illustration generated with Qwen Image.
SHARE
As discussed on the previous post, the analysis of free text on the Web – and as an example the thoughts expressed by Twitter users-  could extract very interesting insights on how users think and how they behave.

In 2001 I visited Trillium where I had a very useful seminar on Data Cleaning, Data Quality and Standardization during which the pareto principle became -once again- evident. When someone wishes to standardize entries in a Database so that the word “Parkway” is written in the same way across all records, he might find the following distribution of “parkway” entries :

15% of records contain the word “Parkway”
3% of records contain the word “Pkwy”
0.2% of records contain the word “Prkwy”
0.01% of records contain the word “Parkwy”

What that essentially means is that with a single SQL query one can find and correct 15% of “parkway” word synonyms to whatever standardized form is needed. But for the remaining variations one query solves only a very small fraction of the problem and this in turn increases the amount of work required, sometimes overwhelmingly.

In capturing and analyzing natural language we are confronted with the same problem : 60% of people might be using the same phrase for describing the fact that they don’t want to go to sleep with a simple “I don’t want to go to sleep”. But another 20% might be using something like : “i don’t feel like sleeping” and another 10% something like “i don’t want to go to bed right now”.

So we immediately see one of the issues that Text Miners face : The fact that we can use different phrases to communicate the same meaning. If we wish to analyze text information for classification purposes -say the sentiment of customers- we could achieve a 60-65% accuracy in our results with some effort. For a mere 4% increase in accuracy -from 65% to 69%- the amount of extra effort required could prove prohibitive.

Consider the following chart :

These are all examples of phrases people use in their everyday talk. We can visualize such phrases starting with” I don’t want to” and then each branch adds a new meaning to the phrase. So branches marked with numbers are the parts of speech that give us an idea of what a person doesn’t want to do : To go, to feel, to visit,to know. Things are getting much more difficult in terms of the effort required if we wish to add more detail -and probably insight- to our analysis by moving further down the branches in our sentence tree.

Perhaps for marketeers, the ability to quantify the distribution of words on the 1st level of the tree depicted above could be enough : If we end up with the following words distribution :

To feel : 15%
To know : 7%
To go : 1%
To visit : 1%

Then, we get an insight on which words to use to market products more efficiently.

On the next post we will go through a hands-on example of analyzing the thoughts of Twitter users and specifically what people seem to “don’t want”.

Share This Article
Facebook Pinterest LinkedIn
Share

Follow us on Facebook

Latest News

How Digital Knowledge Repositories Facilitate Self-Directed Research and Information Discovery -- AI-generated illustration
How Digital Knowledge Repositories Facilitate Self-Directed Research and Information Discovery
Exclusive News
7 MDR Providers Combining Offensive Security Testing With 24/7 Monitoring -- AI-generated illustration
7 MDR Providers Combining Offensive Security Testing With 24/7 Monitoring
Exclusive IT Security
The Information Governance Practices That High-Demand Social Work Roles Require -- AI-generated illustration
The Information Governance Practices That High-Demand Social Work Roles Require
Data Management Exclusive Policy and Governance Security
8 MCP Tools for Market and Consumer Intelligence Workflows -- AI-generated illustration
8 MCP Tools for Market and Consumer Intelligence Workflows
Artificial Intelligence Exclusive

Stay Connected

1.2KFollowersLike
33.7KFollowersFollow
222FollowersPin

You Might also Like

Are Unsubscribe Confirmation Emails CAN-SPAM Compliant?
Uncategorized

Are Unsubscribe Confirmation Emails CAN-SPAM Compliant?

4 Min Read
Schema on Read vs Schema on Write and Why Shakespeare Hates Me
Uncategorized

Schema on Read vs Schema on Write and Why Shakespeare Hates Me

5 Min Read
Social Media Roundup for January 13
Uncategorized

Social Media Roundup for January 13

6 Min Read
Salesforce.com and Oracle: A Tale of Two Worlds
Uncategorized

Salesforce.com and Oracle: A Tale of Two Worlds

5 Min Read

SmartData Collective is one of the largest & trusted community covering technical content about Big Data, BI, Cloud, Analytics, Artificial Intelligence, IoT & more.

5 Great Tips for Using Data Analytics for Website UX
5 Great Tips for Using Data Analytics for Website UX
Big Data
From Bolts to Bots: How AI Is Fortifying the Automotive Industry
From Bolts to Bots: How AI Is Fortifying the Automotive Industry
Artificial Intelligence

Quick Link

  • About
  • Contact
  • Privacy
Follow US
© 2008-26 SmartData Collective. All Rights Reserved.
Welcome Back!

Sign in to your account

Username or Email Address
Password

Lost your password?