Showing posts with label natural language. Show all posts
Showing posts with label natural language. Show all posts

Saturday, August 30, 2025

Benford's Law

Benford’s law describes the relative frequency distribution for leading digits of numbers in datasets. Leading digits with smaller values occur more frequently than larger values. This law states that approximately 30% of numbers start with a 1 while less than 5% start with a 9. According to this law, leading 1s appear 6.5 times as often as leading 9s! Benford’s law is also known as the First Digit Law.

If leading digits 1 – 9 had an equal probability, they’d each occur 11.1% of the time. However, that is not true in many datasets. The graph displays the distribution of leading digits according to Benford’s law.


Analysis of datasets shows that many follow Benford’s law. For example, analysts have found that stock prices, population numbers, death rates, sports statistics, financial and tax information, and billing amounts often have leading digits that follow this distribution. 

Uses for Benford’s Law

Analysts have used it extensively to look for fraud and manipulation in financial records, tax returns, applications, and decision-making documents. They compare the distribution of leading digits in these datasets to Benford’s law. When the leading digits don’t follow the distribution, it’s a red flag for fraud in some datasets.

When Does Benford’s Law Apply and Not Apply

Benford’s law generally applies to data that fit some of the following guidelines:

  • Quantitative data.
  • Data that are measured rather than assigned.
  • Ranges over orders of magnitudes.
  • Not artificially restricted by minimums or maximums.
  • Mixed populations.
  • Larger datasets are better.

Elaborations on Guidelines

Benford’s law often does not apply to assigned numbers, such as ID numbers, phone numbers, and zip codes.

It works best for data that range over multiple orders of magnitudes from very low to very high. You can cover the 10s, 100s, 1000s, and so on. For example, population and incomes can range from very low to very high.

Conversely, if the range of values is restricted, it affects the leading digits, and Benford’s law is less likely to apply. For example, human characteristics naturally fall into restricted ranges. Consequently, this distribution doesn’t apply to human ages, heights and weights. Similarly, limits imposed on potential values can also invalidate this law. Awards in small claims courts have an upper limit, which can negate Benford’s law.

Interestingly, mathematicians have proven that numbers from mixed populations follow Benford’s law. Mixed populations are things like all numbers pulled from a magazine issue. Obviously, those numbers will represent various topics and types of values. Benford himself did that with Reader’s Digest and newspapers. You can also combine data from different sources to achieve the same effect.

Like all distributions, larger datasets will produce observed relative frequencies that more closely approximate the theoretical values of Benford’s law. Smaller datasets can create relatively large deviations due to random error. Some analysts say datasets as small as 100 are acceptable, but most think a minimum size of 500 or even 1,000 is necessary.

Curiously, it will work in some cases where it should not. For example, it applies to house numbers even though those are assigned.

Benford’s Law Formula

Benford’s law formula is the following:

Where d = the values of the leading digits from 1 to 9.

The formula calculates the probability for each leading digit. The table below displays the probabilities that Benford’s law formula calculates for all digits.

Digit     Probability

1     30.1%

2     17.6%

3     12.5%

4     9.7%

5     7.9%

6     6.7%

7     5.8%

8     5.1%

9     4.6%

source: https://statisticsbyjim.com/probability/benfords-law/

Saturday, June 28, 2025

NLWeb and MCP

Natural Language Web

NLWeb, short for Natural Language Web, aims to be the fastest and easiest way to effectively turn your website into an AI app. A natural language interface for websites using the model of their choice and their own data. Every NLWeb instance is also a Model Context Protocol (MCP) server, allowing websites to make their content discoverable and accessible to agents and other participants in the MCP ecosystem if they choose.

How does it work?

NLWeb leverages semi-structured formats like Schema.org, RSS and other data that websites already publish, combining them with LLM-powered tools to create natural language interfaces usable by both humans and AI agents. The NLWeb system enhances this structured data by incorporating external knowledge from the underlying LLMs for richer user experiences.

How do I get started?

The NLWeb GitHub repo contains everything you need to get started:

  • The lightweight code that controls the core service to handle natural language queries, as well as documentation on how this can be extended and customized.
  • Connectors to some of the most popular models and vector databases, as well as documentation to add other models of your choice.
  • Tools for adding your data in Schema.org, JSONL, RSS and other formats to your chosen vector database.
  • A web server frontend for the service and a simple UI that allows users to send queries to the web server.

source: https://news.microsoft.com/source/features/company-news/introducing-nlweb-bringing-conversational-interfaces-directly-to-the-web/

Written by Microsoft Corporate Blogs Published May 19, 2025

Model Context Protocol

The Model Context Protocol, or MCP for short, is a standard for connecting AI assistants to the systems where data resides.

MCP lets AI models draw data from sources like business tools and software to complete tasks, as well as from content repositories and app development environments.

MCP enables developers to build two-way connections between data sources and AI-powered applications (e.g., chatbots). Developers can expose data through “MCP servers” and build “MCP clients” — for instance, apps and workflows — that connect to those servers on command.

source: https://techcrunch.com/2024/11/25/anthropic-proposes-a-way-to-connect-data-to-ai-chatbots/

Transforming the Web with Natural Language: My NLWeb Presentation at Nashua CLOUD .NET & DevBoston

View the slides on SlideShare: https://www.slideshare.net/slideshow/transform-any-website-into-a-conversational-experience-with-nlweb/281034902

What Is NLWeb?

NLWeb (Natural Language Web) is a robust protocol and toolset developed by Microsoft that turns any traditional website into a conversational interface, leveraging the power of large language models. It’s built around the Model Context Protocol (MCP), allowing developers to process natural-language queries and respond using structured Schema.org JSON.

In my session, I demonstrated how NLWeb works, highlighting its design for flexibility (enabling the swapping out of models, vector databases, and embeddings), and how it seamlessly connects to data and APIs to deliver intelligent, real-time responses to users.

Real-World Impact

I also highlighted real-world use cases where NLWeb is already in action:

  • Tripadvisor – enabling users to plan family trips through conversation
  • Eventbrite – allowing event discovery through natural-language search
  • O’Reilly, Qdrant, Delish, Shopify, and others – showcasing early success in turning structured content into AI-driven UX

These examples demonstrate how businesses are already leveraging the potential of conversational web interfaces to drive engagement and discovery.

How to Get Started

For those interested in experimenting or building with NLWeb, here are a few resources I shared:

  • GitHub: https://github.com/microsoft/NLWeb
  • Quick start guide: docs/nlweb-hello-world.md
  • Local test interface: http://localhost:8000/static/debug.html
  • Azure deployment: docs/setup-azure.md

Whether you’re a developer, architect, or product leader, NLWeb offers a modern and modular approach to embedding LLM-driven intelligence into any web property.

source: https://udai.io/transforming-the-web-with-natural-language-my-nlweb-presentation-at-nashua-cloud-net-devboston/

Saturday, January 11, 2025

O.S.A.S.C.O.M.P. (adjective order in English)

"The rule is that multiple adjectives are always ranked accordingly: opinion, size, age, shape, colour, origin, material, purpose. Unlike many laws of grammar or syntax, this one is virtually inviolable, even in informal speech. You simply can’t say My Greek Fat Big Wedding, or leather walking brown boots." ~ Tim Dowling in The Guardian (full article)

source: https://sketchplanations.com/ordering-adjectives

But of course, it's not that simple. This list and its order vary. According to the Cambridge Dictionary, the order contains 10, not 8, categories, and is:

order

relating to

examples

1

opinion

unusual, lovely, beautiful

2

size

big, small, tall

3

physical quality

thin, rough, untidy

4

shape

round, square, rectangular

5

age

young, old, youthful

6

colour

blue, red, pink

7

origin

Dutch, Japanese, Turkish

8

material

metal, wood, plastic

9

type

general-purpose, four-sided, U-shaped

10

purpose

cleaning, hammering, cooking

Some prefix the list with determiner and quantity. source: https://byjus.com/english/order-of-adjectives

This English as a second language (ESL) site, eslgrammar.org, is more forgiving about the exact order by grouping size, shape, age, and color together.

Hope you enjoyed reading this English short surprising article. Or should it be surprising short English article?

Saturday, September 14, 2024

UPPERCASE is better than lowercase when comparing text

If locale-based string comparison is not available, converting to UPPERCASE is better than lowercase. For example,

uppercase("Straße") == uppercase("Strasse") == "STRASSE"

lowercase("Straße") != lowercase("Strasse"), because "straße" != "strasse"

For more on the ß character, see Eszett (ß = German sharp S).

UPDATE: According to Wikipedia, "Traditionally, ⟨ß⟩ did not have a capital form, although some type designers introduced de facto capitalized variants. In 2017, the Council for German Orthography officially adopted a capital, ⟨ẞ⟩, as an acceptable variant in German orthography, ending a long orthographic debate. Since 2024 the capital ⟨ẞ⟩ (ligature) has been preferred over ⟨SS⟩ (two letters)."

If you are making a decision based on the result of a string comparison or a case change, be sure to avoid the invariant culture. See Invariant Culture.

Saturday, July 13, 2024

The Mandela Effect

The Mandela Effect refers to widespread false memories that large numbers of people or a group of individuals believe. Memory is not a perfect recording of events that happened. It can change with time and with practice and priming.

source: Mandela Effect: Examples and explanation - MedicalNewsToday

The term was originated in 2009 by Fiona Broome, after she discovered that she, along with a number of others, believed that Nelson Mandela had died in the 1980s (when he actually died in 2013).  

Some explanations

  • False Memories: similar memories are associated with each other
  • Confabulation: filling in gaps that are missing to make more sense of them
  • Misleading Post-Event Information
  • Priming: factors leading up to an event that affects perception

source: Mandela Effect Examples, Origins, and Explanations   

More examples

 

Saturday, July 6, 2024

logophobia, verbophobia

Meaning: Fear of words

This fear typically originates from childhood, where the frequency of learning new words can cause distress and dread. Another cause is frustration from frequent misspellings, such as might occur in a spelling bee.

For me, the problem is reading misspelled words, such as in newspaper articles. Being able to spot defects is an asset in software engineering, but hampers dealing with natural language. Reading poorly written content is slow and frustrating. Writing, likewise, is slow and frustrating in an effort to avoid creating poorly written content.

sources:

  • https://phobia.fandom.com/wiki/Logophobia
  • https://en.wiktionary.org/wiki/logophobia
  • https://en.wiktionary.org/wiki/verbophobia


Saturday, June 29, 2024

glossophobia

 Meaning: Fear of public speaking

The word glossophobia derives from the Greek γλῶσσα glossa (tongue) and φόβος phobos (fear or dread.)

source: https://en.wikipedia.org/wiki/Glossophobia

Saturday, June 22, 2024

sesquipedalophobia

Meaning: Fear of long words

Root from Latin sēsquipedālis (literally “a foot and a half long”), from Latin sēsqui (“one and a half times”) + Latin pedālis (“measuring a foot, foot (relational)”)

And to make it even longer, non-nonsensical parts were added to make: 

hippopotomonstrosesquipedaliophobia

sources: 

  • https://en.wiktionary.org/wiki/sesquipedalophobia#English
  • https://en.wiktionary.org/wiki/sesquipedalian#English 
  • https://en.wiktionary.org/wiki/hippopotomonstrosesquipedaliophobia

Saturday, April 30, 2022

Weird

i before e except after c
or when sounded as 'a' as in neighbor and weigh

Nice try, but it's still full of exceptions. To make the above jingle accurate, it'd need to be something like:

I before e, except after c
Or when sounded as 'a' as in 'neighbor' and 'weigh'
Unless the 'c' is part of a 'sh' sound as in 'glacier'
Or it appears in comparatives and superlatives like 'fancier'
And also except when the vowels are sounded as 'e' as in 'seize'
Or 'i' as in 'height'
Or also in '-ing' inflections ending in '-e' as in 'cueing'
Or in compound words as in 'albeit'
Or occasionally in technical words with strong etymological links to their parent languages as in 'cuneiform'
Or in other numerous and random exceptions such as 'science', 'forfeit', and 'weird'.

source: https://www.merriam-webster.com/words-at-play/i-before-e-except-after-c

Weird has a weird spelling. It's not only fails the rule, it fails the exceptions.




Friday, October 16, 2020

numeronym

A numeronym is a number-based word.

Most commonly, a numeronym is a word where a number is used to form an abbreviation. Pronouncing the letters and numbers may sound similar to the full word: "K9" for "canine".

Alternatively, the letters between the first and last are replaced with a number representing the number of letters omitted, such as "i18n" for "internationalization". These word shortenings are sometimes called alphanumeric acronyms, alphanumeric abbreviations, or numerical contractions.

According to Tex Texin, the first numeronym of this kind was "S12n", the electronic mail account name given to Jan Scherpenhuizen by a system administrator because his surname was too long to be an account name. 

A number may also denote how many times the character before or after it is repeated. This is typically used to represent a name or phrase in which several consecutive words start with the same letter, as in  W3C (World Wide Web Consortium).

Examples

g11n – globalisation / globalization
a11y – accessibility
tr8n – translation
l10n – localisation / localization
i18n – internationalization
k8s – Kubernetes

Sunday, June 9, 2019

negativity bias


Daniel Goleman research shows what happens when people talk face to face or on the phone. The connections that their brains make, and how this (usually) allows for some form of responsiveness within the conversation depending on the reaction of the other person. This does not work with emails, texts or online communication.

Goleman has found that there is a negativity bias online. That is “..what you thought was a neutral message can be perceived as hostile by the recipient..” and what you thought was positive can be perceived as neutral by the recipient.

source: Combating negative bias in communication

Friday, December 21, 2018

Stop words

In computing, stop words are words which are filtered out before or after processing of natural language text. For some search engines, these are some of the most common, short function words, such as the, is, at, which, and on.

A predecessor concept was used in creating some concordances. For example, the first Hebrew concordance contained a one-page list of unindexed words

source: https://en.wikipedia.org/wiki/Stop_words

Saturday, September 8, 2018

n-gram model

In the fields of computational linguistics and probability, an n-gram is a contiguous sequence of n items from a given sample of text or speech. The items can be phonemes, syllables, letters, words or base pairs according to the application. The n-grams typically are collected from a text or speech corpus. When the items are words, n-grams may also be called shingles. Using Latin numerical prefixes, an n-gram of size 1 is referred to as a "unigram"; size 2 is a "bigram" (or, less commonly, a "digram").

An n-gram model is a type of probabilistic language model for predicting the next item in such a sequence in the form of a (n − 1)–order Markov model. n-gram models are now widely used in probability, communication theory, computational linguistics (for instance, statistical natural language processing).

n-grams can also be used for efficient approximate matching. n-grams have been used to design kernels that allow machine learning algorithms such as support vector machines* to learn from string data.

To choose a value for n in an n-gram model, it is necessary to find the right trade off between the stability of the estimate against its appropriateness. This means that trigram (i.e. triplets of words) is a common choice with large training corpora (millions of words), whereas a bigram is often used with smaller ones.

source: https://en.wikipedia.org/wiki/N-gram

BONUS: Google Ngram Viewer - historical frequency of some AI terms
 

* In machine learning, support vector machines (SVMs, also support vector networks) are supervised learning models with associated learning algorithms that analyze data used for classification and regression analysis.

source: https://en.wikipedia.org/wiki/Support_vector_machine

TF-IDF term frequency / inverse document frequency

TF-IDF stands for “term frequency / inverse document frequency” and is a method for emphasizing words that occur frequently in a given document, while at the same time de-emphasising words that occur frequently in many documents.

source: http://fastml.com/classifying-text-with-bag-of-words-a-tutorial/

Bag-of-words model

The bag-of-words model is a simplifying representation used in natural language processing and information retrieval (IR). The bag-of-words model is commonly used in methods of document classification where the frequency of occurrence of each word is used as a feature for training a classifier.

In this model, a text (such as a sentence or a document) is represented as the bag (multiset) of its words. Each key is the word, and each value is the number of occurrences of that word in the given text document.

Example usage: spam filtering

In Bayesian spam filtering, an e-mail message is modeled as an unordered collection of words selected from one of two probability distributions: one representing spam and one representing legitimate e-mail. To classify an e-mail message, the Bayesian spam filter assumes that the message is a pile of words that has been poured out randomly from one of the two bags, and uses Bayesian probability to determine which bag it is more likely to be.

source: https://en.wikipedia.org/wiki/Bag-of-words_model

Sunday, August 12, 2018

compile

compile

from Latin compilare "to plunder, rob," from com- "together" + pilare "to compress, ram down." ~ Online Etymology Dictionary

Eszett (ß = German sharp S)

Eszett (ß = German sharp S)
In the German alphabet, ß (Unicode U+00DF) is a consonant letter that evolved as a ligature of "long s and z" (ſz) and "long s over round s" (ſs).
The combination of long s and s is also seen in Early Modern English (example from the US Bill of Rights)
Since the German spelling reform of 1996, both ß and ss are used to represent /s/ between two vowels. In alphabetizing German words, the collation rules say to treat ß as a double "s". 

Source: http://en.wikipedia.org/wiki/%C3%9F

If locale-based string comparison is not available, converting to UPPERCASE is better than lowercase. For example,

uppercase("Straße") == uppercase("Strasse") == "STRASSE"

lowercase("Straße") != lowercase("Strasse"), because "straße" != "strasse"

UPDATE: According to Wikipedia, "Traditionally, ⟨ß⟩ did not have a capital form, although some type designers introduced de facto capitalized variants. In 2017, the Council for German Orthography officially adopted a capital, ⟨ẞ⟩, as an acceptable variant in German orthography, ending a long orthographic debate. Since 2024 the capital ⟨ẞ⟩ (ligature) has been preferred over ⟨SS⟩ (two letters)."

Engineer on WordMap





Many of the words have audio pronunciations. Click on a word to hear it spoken.

The Indo-European connection is easily seen (click on the word in India).

The English word ‘mechanic’ is a synonym (click on Greece) and the Arabic-related languages have a word with some similarity to ‘mechanic’: muhundis.