Showing posts with label Bayes' Theorem. Show all posts
Showing posts with label Bayes' Theorem. Show all posts

Tuesday, April 16, 2013

Is It Safe To Go Out Yet? Statistical Inference in a Zombie Outbreak Model

Summary: 
The authors Calderhead, Girolami, and Higham (2010) wrote a paper dealing with the potential outcomes of a zombie outbreak using Bayesian theorem to support their conclusions.  Since there has never been a zombie outbreak it is logical to use Bayesian theory to account for unknown data that can be estimated.  Estimations can be made regarding a zombie outbreak, and in turn these estimations can be used in particular ways to end with likely outcomes.

The idea for using Bayesian theory applied to zombie outbreaks starts with a logical probability.  In the case of this paper, the authors state that in one day the far extremes of probability are that no human will turn into a zombie and all humans are converted into zombies.  This probability (labeled prior) is then updated as new data is used (such as different quantity of days) and thus this posterior distribution then becomes the prior and the process is repeated.

The authors then state that many questions can be answered by successfully finding a likely distribution for human to zombie conversion rates.  Such questions include how many soldiers should be mobilized, the scale of quarantine needed, and whether or not it is alright to leave a hiding spot given the number of zombie sightings during a particular time span.  The authors also emphasize that since the rate of change from human to zombie is likely to not be constant, the beta (conversion coefficient) should be a range and not a singular number.

One of the model comparisons the authors use is the comparison of two models: one that assumes zombies can attack alone while the other shows that after a rumor is circulated, zombies are believed to only travel in pairs.  The authors then seek to disprove the second model through Bayes factors (posterior odds = Bayes factor * prior odds), in which statistical evidence for the first model is weighed heavier than the second.  The authors find that the first model with the least amount of noise (introduced as Gaussian distributed noise) is more likely.  This means that the experimental data deviated the least from the expected curve.  Adding additional noise negatively impacted the Bayes factor (shows as how strong the evidence was against the second model.


Another example the authors used Bayesian theory on was for answering the question of whether or not it is safer to leave a hiding spot based on the number of zombies spotted in the past few days.  The authors use two types of analysis for this: one that does not have any observations of previous day's zombie sightings and one that has five day totals of zombie sightings.  The stated the zombie sightings for the second case were 123, 127, 104, 92, and 74.  The left column of the figure below shows Bayesian factors applied to the first model mentioned above, and the potential outcomes.  The right column shows this process with a second layer of data (the five days of observations) which greatly reduces the uncertainty regarding potential zombie totals for the next 45 days.  Thus, due to these observations and using Bayesian theory uncertainty can be greatly reduced and the chance of surviving longer during a zombie outbreak are much higher.



Critique:
I found this article to be fairly complex to read with no experience with Bayesian theory.  However, I really appreciate the application of Bayesian theory to a zombie outbreak.  Although this article topic is (most likely) fantastical, it is well constructed and thoughtful.  Honestly, the topic caught my eye and I doubt I would have tried as hard as I did to understand Bayesian theory had it been on a drier topic.

One issue I had with this article is that it was clearly meant for someone that had previous experience with Bayesian theory.  At times the authors referenced various aspects of Bayesian theory without defining them.  For example, the authors did not just list Bayes factors but instead referenced an article.  For the average reader this does not make reading and understanding this topic any easier.  Additionally, simple Bayesian theory is not that long or difficult to write out and would have saved me time having to look it up to double check I was thinking about the correct thing.

This article is not related to intelligence, save for if there were to ever be a zombie attack it would prove useful for intelligence analysts to extend their lifespans.  However, this model could be applied to medicine for spread of infectious diseases if the transmission rate is unknown.  Instead of zombies and humans there would be infected and healthy individuals.

Source:
Calderhead, B., Girolami, M., & Higham, D. (2010). Is It Safe To Go Out Yet? Statistical Inferencein a Zombie Outbreak Model.University of Strathclyde, United Kingdom.  Retrieved from http://www.strath.ac.uk/media/departments/mathematics/researchreports/2010/6zombierep.pdf

Tuesday, April 17, 2012

Applying Bayesian technique to filtering spam

Introduction
The white paper, "Why Bayesian filtering is the most effective anti-spam technology" describes how a company can apply Bayesian mathematics to spam e-mails’ problem by creating an adaptive, ‘statistical intelligence’ technique that increases spam detection rates. The unique characteristic that distinguishes from other spam filters is that the company or organization can customize the filter based on company or organization’s email characteristics, and update the database with newly detected spam characteristics.

Summary
Bayesian filtering is based on the principle that most events are dependent and that the probability of an event occurring in the future can infer from previous occurrences of that event. According to the paper, before we can filter span using this method, we need to generate a database with words and token collected from a sample of spam mail and valid mail, also referred to as ‘ham’. A probability value is assigned to each word or token, which is based on calculations that take into account how often that word occurs in spam as opposed to legitimate mail (ham). The word probability is calculated by analyzing the users’ outbound mail and by analyzing known spam.

For example: If the word “mortgage” occurs in 400 of 3,000 spam mails and in 5 out of 300 legitimate emails, then its spam probability would be { [400/3000] divided by [(5/300) + (400/3000)] } i.e. 0.8889.

Creating the ham database: The analysis of ham mail is performed on the organization’s mail, and is therefore tailored to that particular organization. A financial institution, for example, might use the word “mortgage” many times over and would get a lot of false positives if using a general anti-spam rule set. However, the Bayesian filter, if tailored to your company through an initial training period, takes note of the company’s valid outbound mail and recognizes “mortgage” as bring frequently used in legitimate messages, and therefore has a much better spam detection rate and a far lower false positive rate.
Creating a spam database: Along with ham database, the Bayesian filter also relies on spam data file. The spam data file must include a large sample of known spam and must be constantly updated with the latest spam by the anti-spam software. This will ensure that the Bayesian filter is aware of the latest spam tricks, resulting in a high spam detection rate.
How the actual filtering is done: Once the ham and spam databases have been created, the word probabilities can be calculated and the filter is ready for use. When a new mail arrives, it is broken down into words and the most relevant words, such as those that are most significant in identifying whether the mail is spam or not, are singled out. From these words, Bayesian filter calculates the probability of the new message being spam or not. If the probability is greater than a threshold, 0.9 for instance, then the message is classified as spam. This Bayesian approach to spam is highly effective, a May 2003 BBC article reported that spam detection rates of over 99.7 percent can be achieved with a very low number of false positives.

Conclusion
Bayesian filtering, if implemented the right way and tailored to an organization, is the most effective technology to combat spam. The downside to this technique is that, we have to wait at least two weeks for it to learn and create the ham or spam database. Nevertheless, over time, the Bayesian filter becomes more and more effective as it learns about the organization’s email habits, along with updating through other anti-spam databases. 

Source:
White Paper, GFI (2008). Why Bayesian filtering is the most effective anti-spam technology. Retrieved from http://www.gfi.com/whitepapers/why-bayesian-filtering.pdf

Monday, April 16, 2012

Measuring Sustained Competitive Advantage Using Bayesian Reasoning

Introduction:
In this article, Tang and Liou develop a theoretical framework to understand the causal relationships among (1) sustainable competitive advantage, (2) configuration, (3) dynamic capability, and (4) sustainable superior performance. They propose that a firm’s competitive advantage, resource bundle configuration, and dynamic learning capability cannot be comprehended by outsiders. Its operational performance, however, can be captured by financial indicators.

They promote an inductive Bayesian interpretation of the sustainable competitive advantage proposition. From this viewpoint, the presence or absence of competitive advantage may be reflected in the causal relationship between resource configuration, dynamic capability, and observable financial performance. They then apply this theoretical framework to an example drawn from the global semiconductor industry, an area in which resource configuration and dynamic capability are essential to performance.

Summary:
The paper expounds on the theory supporting the use of Bayesian reasoning to measure competitive advantage over other measures like Porter’s competitive strategy or the resource based view of valuable, rare, inimitable and non-substitutable.

Powell’s premise – sustainable competitive advantage is more probable in firms that have already achieved sustained superior performance – is further developed through periodically updating its propositions or hypotheses in the face of empirical evidence. Tang and Liou use a Bayesian discriminate model to reveal the functional dependence of superior performance on unique business processes. The primary sources of competitive advantage are considered embedded in and inseparable from the organization itself, along with its business units and functional departments. It is assumed that the process of managing these resources, termed strategic fit, cannot be comprehended or imitated by outsiders. The model is then demonstrated using the semi-conductor industry.
Explanation of sustainable competitive advantage
Model:
The final data set for the study contained 147 companies and 786 firm-year observations. Of those, 188 companies are located in developed countries (The U.S., within Europe, and Japan). The other 29 are in the Asia/Pacific region. Using the firm’s financial data, certain observable traits can be inferred.

The study began by using PCA, principle component analysis, on the financial indicators to identify the traits or factors. Three principal factors accounted for 60 percent of the total variance.

Factor 1: Relationship management. This factor includes customer relationship management (accounts receivable turnover), three variables related to supplier relationship management (accounts payable turnover, inventory turnover, and CGS/sales) and one variable associated with the government (tax to sales ratio). The factor illustrates the sustainable competitive advantage of firms that manage upstream (suppliers), downstream (customers), and governmental relationships. The variance indicated that good relationship management can pay off with respect to a lower CGS.

Factor 2: Management ability. This factor consists of indicators related to fixed asset management capabalities including department/sales ratio and fixed asset turnover. The correlation between fixed assets turnover and Factor 2 indicates that firms with greater competence in assets management generate revenue at a lower unit cost and low asset depreciation.

Factor 3: Knowledge management. This factor includes R&D/sales and SG&A/sales ratios to measure a firm’s effectiveness in resource deployment. The high correlation indicates that lower unit costs are associated with efficient management.

The findings, therefore, support the idea that resource configurations or factors of a firm can be inferred from their observable financial indicators.

Conclusion:
Tang and Liou advance Powell’s idea of using Bayesain probabilistic reasoning as a means of distinguishing sustained competitive advantage from sustained superior performance in this paper.
They propose that particular resource configurations can be shown to link the two – sustained competitive advantage and sustained superior performance.

Through a discussion of Bayes’ theory and subsequent semi-conductor example, the paper describes how empirical data on past financial performance in a population of firms can be used to generate the posterior probability of sustainable competitive advantage, given the prior probabilities of both competitive advantage and competitive disadvantage.

Source:Tan, Y-C; Liou, F-M. (2010). Does Firm Performance Reveal Its Own Causes? The Role of Bayesian Inference. Strategic Management Journal. 31: 39-57.
Retrieved from http://web.it.nctu.edu.tw/~etang/SMJ2010_TangLiou.pdf

Likelihood of Global Warming Given X.

Introduction:

I’m going to profile a rather wacky article here. A Norwegian physics student applied Bayes’ Theorem to a number of things in a peripheral way in order to provide a basis for applying the theory to global warming. In “Testing Hypotheses about Climate Change: the Bayesian Approach”, Kristoff Rypdal applies Bayes to Russian roulette, the learning process of individuals, hurricanes and global warming and the melting of arctic ice caps. Essentially the author applied probability theory and hypothesis testing, where the concept of probability is defined subjectively as “a degree of knowledge" about a hypothesis. He defines knowledge as something generated by four processes: 1) the inspired formulation of new hypotheses 2) prediction (here deductio

n enters a central element) 3) collection of new data through experiment or observation 4) verification/falsification by comparing predictions and observations.

Summary:

I’ll go ahead and skip Kristoff’s lengthy explanation of what Bayes’ theory does via a parable about Mafiosos and their desire to watch him play Russian roulette. After by passing his .38, Kristoff takes us to Section VI. Hurricanes and Global Warming. He essential states that for the sake of the formula, the existence of human caused global warming is bivalent, 50/50 (H,HN). The scientific community (again for the sake of the formula) thinks that the odds of a massive hurricane occurring more than once per century in the absence of global warming is 10%, i.e. p(B|HN)=0.1. He also states that the scientific community believed that in the presence of global warming, massive hurricanes will occur more than once per century is 50%, i.e. p(B|H)=0.5.


Kristoff then basically applies the exact same theory to the arctic ice caps, in Section VIII: Bayesian Learning and Arctic Ice Cap Melting. He states that there has been a well-documented scientific effort regarding the monitoring of the arctic ice caps (which is true) and that these scientists have observed a large reduction in summer sea ice (which is also true). He then gives the “scientific estimates of probability” for the sake of applying the situation to Bayes’ Theorem. He states that the probability of the arctic ice caps melting w/o human caused global warming is 10% or, p(C|HN)=.1. He also states that the odds of this kind of melting occurring in conjunction with the presence of human caused global warming is 50% or, p(C|H)=0.5. These two measurements are identical as the one above and result in p(H|C)=.83 or 83%. When combined however, the author states that p(HN)=1-p(H)=.17. When combining this to p(H|B) we get .96, thus turning the results of the previous equation from ‘highly likely’ to ‘virtually certain’.

Conclusion:

Overall I thought Kristoff’s application was interesting though only theoretical in nature. It would be interesting to conduct polling studies on the topic of global warming in both the scientific community and the general public and apply Bayes theorem to the results. This could actually be applied to any polling of public perception really, so long as there was actual information to prove something correct or incorrect. As a study of ‘the causes of global warming’ Kristoff’s article is speculative and not particularly edifying, but as a study of human perception and likely, it is rather interesting.

Source:

Rypdal, K. (2008). Testing hypotheses about climate change: the Bayesian approach. Department of Physics and Technology, University of Troms, 9037 Troms, Norway

http://web.me.com/kristofferrypdal/Themes_Site/Courses_files/Bayesian%20approach%20to%20Climate%20Change.pdf

Wednesday, May 6, 2009

Summary Of Findings: Bayesian Analysis (4 out of 5 Stars)

Note: This post represents the synthesis of the thoughts, procedures and experiences of others as represented in the 12 articles read in advance of (see previous posts) and the discussion among the students and instructor during the Advanced Analytic Techniques class at Mercyhurst College on 6 MAY 2009 regarding Bayesian Analysis specifically. This technique was evaluated based on its overall validity, simplicity, flexibility and its ability to effectively use unstructured data.

Description:
Bayesian analysis is a method that uses Bayesian statistics to assess the likelihood of an event happening in light of new evidence. It generates an estimate and the use of Bayesian statistics in Intelligence analysis allows for the uncertainty of the traditional intelligence data set to be understood in a scientifically valid manner.

Strengths:
*can limit analyst biases by reducing the weight of evidence simply because it is new or vivid
*forces the analyst to resassess evidence and consider alternative possibilites
*adheres to rigid mathematical formulas
*provides a numerical likelihood
*provides audit trail and ability to reproduce results

Weaknesses:
*Probabilities are based largely on subjectivity
*Susceptible to biases
*Highly complex problems require heavy computations
*Can be mathematically complex
*Not always useful as a stand alone method (works well in tandem with methods like Delphi); may require SMEs for determining probability distributions
*Some reliance on ambiguous validities
*"Negative evidence"--absence of positive evidence

How-To:

This method loosely follows the guidance suggested by his line of research into the use of natural frequencies in teaching and explaining Bayes to beginners.

1.) Create a 2x2 matrix. Label the quadrants with the respective information that creates true positive, false negative, false positive, and true negative quadrants.
2.) Take the given information, the base line (for example, 100 out of 1,000) with the new information (for example, a new document that is 90% credible saying that war is immiment) which means that your true positive and your false negative must equal 100 and the false positive and true negative must equal 900.
3.) To calculate the true positive quandrant, take 90% of the 100 from the base line (which equals 90).
4.) To calculate the false negative quadrant, take the numerator of the base line (100) and subtract the true positive quadrant (90), creating an answer of 10.
5.) To calculate the true negative quadrant, take 90% of your non-war cases (900), equalling 810.
6.) To calculate the false positives, subtract the sum of the three quadrants known from the total number of cases (1,000), which equals 90.
7.) To calculate the new probablitiy, divide what the numerator of the base line (100) from the new total of positive caes (90+90=180), which equals 55.5%

The 55.5% means that there is a 55% probability that countries X and Y are likely to go to war.



Experience:
To understand the basic mathematical principles behind Bayes, the class worked through some sample problems. One of the problems was based on a medical test with an 80% accuracy rate for a cancer with 2% affliction rate in the general population. The class applied this to a sample population of 1000 cases. We established a matrix and assessed the true positive, false positive, false negative, and true negative quantities (16, 116, 4, and 784 respectively). We plugged these numbers into the appropriate matrix fields. We then divided the number of actual cases of cancer (20 or 2% of 1000) into the number of positive tests (132--the 16 true positives and 116 false postives). The result was 15% rate of those who have the cancer from the positive tests, a rather stark difference from the 2% base rate! This problem actually reflects the number of breast cancer rates from a medical treatment from around a two decades ago!
Note: see the matrix for a synopsis of another of the problems we worked through (a peice of evidence emerging suggesting a cause for war).

The class also used a Bayesian application to assess the likelihood we would contract swine flu. We started with the initial hypothesis that we would contract swine flu or we would not contract swine flu, and assigned an initial probability to each hypothesis (the latter >5%) We then added weighted evidence which influenced the base rate of the hypothesis. After all the evidence was entered, the class assessed the likelihood of contracting swine flu.

Saturday, May 2, 2009

Bayes' Theorem and Intelligence

Net Wars

Summary:
According to the author, "Beliefs are based on probabilistic information. Bayes Theorem says that our initial beliefs are updated to to posterior beliefs after observing new conditions." As an analytic method, Bayesian analysis provides a formula which allows the analyst to upgrade original assertions as new evidence is discovered, and assign likelihoods to events; the more we observe the better we can predict the likelihood of a certain event. The formula used to update the analysts initial beliefs to posterior beliefs is: p(C|O) = p(O|C)p(C)/p(O|C)p(C) + p(O|¬C)p(¬C). According to Bayes, initial beliefs have a high margin of error; this is alleviated by incorporating new evidence through this formula. This formula "produces interesting results because it accounts for uncertainties created by False Positives and False Negatives."

The author provides the following example of using Bayesian analysis to update your beliefs:

"There is a case of this occurring. Europeans believed that swans were always White and there could be no Black Swans. They updated their probability of a Swan being White to 99% based on their limited experiences. As they explored the world, they found Black Swans in Australia. This reduced the probability of a swan being white and increased the probability of a swan being black. This process of inductive reasoning can be explained via Bayesian probability."

The author provides the following example of using Bayesian analysis as an intelligence methodology:

"There are 10,000 civilians. 1% of whom are insurgents pretending to be civilians. Police can investigate individuals and determine if they are an insurgent or civilian with 95% certainty.

Prior Probability is this: 0.01 (10,000) and 0.99(10,000). So
Group 1: 100 insurgents
Group 2: 9,900 Civilians

The Police investigate the entire population. This produces four groups:
Group 1: Insurgents - Positive test (0.95)
Group 2: Insurgents - False Negative test (0.05)
Group 3: Civilians - False Postive test (0.05)
Group 4: Civilians - Negative test (0.95)

How certain are the police that the men they captured are actually insurgents? The answer is 16%.
(0.95 x 0.01)/ (0.95 x 0.01) + (0.05 x 0.99) =
0.0095/0.0590 = 0.161"

The 16% certainty rate stems from the uncertainty that always exists as some insurgents escape detection while some innocents test positive as insurgents. The author notes that this is an extremely oversimplified example; actual Bayesian analysis in this situation would require some serious computing power that takes into account many other factors, as well as multiple testing to insure that the most accurate results were reached. Nonetheless, the example highlights the use of Bayesian analysis as a method of predictive analysis. However, the method predicts the probability of a particular event happening, and not whether that event will actually occur.

The author returns to the "black swan" example, stating that these highly unlikely, yet possible, intelligence "black swans" are events that can occur, but are highly unlikely too. Just because they haven't happened, doesn't mean they won't. Bayesian analysis provides a method for determining their likelihood. The author concludes by reiterating that intelligence analysis is not about predicting future events, but rather about predicting the likelihood of future events. "The inability to stop a Black Swan event, or a false prediction of a Black Swan event, does not always mean that the intelligence community 'failed'." Rather, the notion that Intel failed comes from the distorted view of Intel analysts as fortune tellers, rather than the reducers of uncertainty that they truly are.

Wednesday, April 29, 2009

Bayes' Formula



Author's Note: This is a great video for teaching Bayes' Theorem in its simplest form.


Summary
In order to illustrate the utility of Bayes’ Theorem, the author draws upon two simple scenarios. First, suppose someone faces the decision of needing to choose between three doors. If the person making the decision does not have any prior knowledge about the situation, the scenario creates an unconditional probability. But, once the person receives new information about the scenario, the rational person should reconsider his/her decision and subsequent probabilities.

Bayes’ Theorem is about the introduction of new information used to adjust probabilities and create conditional probabilities. In the formula, P(G/U), P is the probability that G will occur, if U happens.

To illustrate the application of Bayes’ Theorem and conditional probabilities, the author illustrates a second scenario. Pretend that there is a 70% probability that the economy will grow and a 30% probability that the economy will slow (an unconditional probability). The author owns a stock that has an 80% chance of increasing if the economy grows. That same stock, however, only has a 30% chance of increasing if the economy slows. The 80% and the second 30% are conditional probabilities; they are based on the condition that the economy will grow or slow.

The author can then determine the scenario’s four conditional probabilities:
1) What is the probability that the economy will grow and the stock will increase?
2) What is the probability that the economy will grow and the stock will decrease?
3) What is the probability that the economy will slow and the stock will increase?
4) What is the probability that the economy will slow and the stock will decrease?

To answer these questions and determine their probabilities, the author uses the equation: P(UG) = P(U/G)P(G). Notice that this equation is longer than the first because this one incorporates two conditions: the economy will grow/slow and the stock will increase/decrease.
Reblog this post [with Zemanta]

Bayes' Theorem for Intelligence Analysis

Jack Zlotnick
CIA Historical Review Program


Author’s Note: Released by the CIA’s Historical Program in the early 1990’s, Jack Zlotnick wrote this piece in the 1970’s. At the time, the CIA was still in the process of testing Bayes’ Theorem. Due to the ongoing testing period (at that time), Zlotnick does not offer a position on the utility and validity of the Bayesian method with regards to intelligence. In fact, Zlotnick spends a considerable amount of time in the article discussing the ways the theorem should continue to be tested.

Summary
Due to the very nature of intelligence, analysts should be naturally interested in the Bayesian Theorem. Intelligence is probabilistic in nature. Intelligence analysts usually conduct their analysis based on incomplete evidence in which they must address probabilities (thus WOEP’s).

For intelligence applications, Bayes’ Theorem is represented by the equation R=PL. “R” is the revised estimate of the odds favoring one hypothesis over another competing hypothesis (the odds of a particular hypothesis occurring after new evidence is entered into the equation). “P” is the prior estimate on the hypotheses probabilities (the odds before considering the new evidence entered into the equation). The analyst must offer judgments about “L” or the likelihood ratio. This variable is the analyst’s evaluation of the “diagnosticity” of an item of evidence. For instance, if a foreign power mobilizes its troops, what are the chances that “X” will happen over “Y”.

The principle features of the Bayes Theorem distinguish it from conventional intelligence analysis in three ways. First, it forces analysts to quantify judgments that are not ordinarily expressed in numeric terms. Second, the analyst does not take the available evidence as given and draw conclusions. And third, the analyst makes his/her own judgments about the bits and pieces of evidence. He/she does not sum up the evidence as he/she would if he/she had to judge its meaning for a final conclusion. The mathematics does the summing up.

The author is skeptical that the complex tasks analysts are forced to consider can be reduced to numeric values. Bayes’ Theorem, however, may be useful for examining strategic warning by uncovering patterns of activity by foreign powers.
Reblog this post [with Zemanta]