Showing posts with label Bayesian probability. Show all posts
Showing posts with label Bayesian probability. Show all posts

Tuesday, April 17, 2012

Applying Bayesian technique to filtering spam

Introduction
The white paper, "Why Bayesian filtering is the most effective anti-spam technology" describes how a company can apply Bayesian mathematics to spam e-mails’ problem by creating an adaptive, ‘statistical intelligence’ technique that increases spam detection rates. The unique characteristic that distinguishes from other spam filters is that the company or organization can customize the filter based on company or organization’s email characteristics, and update the database with newly detected spam characteristics.

Summary
Bayesian filtering is based on the principle that most events are dependent and that the probability of an event occurring in the future can infer from previous occurrences of that event. According to the paper, before we can filter span using this method, we need to generate a database with words and token collected from a sample of spam mail and valid mail, also referred to as ‘ham’. A probability value is assigned to each word or token, which is based on calculations that take into account how often that word occurs in spam as opposed to legitimate mail (ham). The word probability is calculated by analyzing the users’ outbound mail and by analyzing known spam.

For example: If the word “mortgage” occurs in 400 of 3,000 spam mails and in 5 out of 300 legitimate emails, then its spam probability would be { [400/3000] divided by [(5/300) + (400/3000)] } i.e. 0.8889.

Creating the ham database: The analysis of ham mail is performed on the organization’s mail, and is therefore tailored to that particular organization. A financial institution, for example, might use the word “mortgage” many times over and would get a lot of false positives if using a general anti-spam rule set. However, the Bayesian filter, if tailored to your company through an initial training period, takes note of the company’s valid outbound mail and recognizes “mortgage” as bring frequently used in legitimate messages, and therefore has a much better spam detection rate and a far lower false positive rate.
Creating a spam database: Along with ham database, the Bayesian filter also relies on spam data file. The spam data file must include a large sample of known spam and must be constantly updated with the latest spam by the anti-spam software. This will ensure that the Bayesian filter is aware of the latest spam tricks, resulting in a high spam detection rate.
How the actual filtering is done: Once the ham and spam databases have been created, the word probabilities can be calculated and the filter is ready for use. When a new mail arrives, it is broken down into words and the most relevant words, such as those that are most significant in identifying whether the mail is spam or not, are singled out. From these words, Bayesian filter calculates the probability of the new message being spam or not. If the probability is greater than a threshold, 0.9 for instance, then the message is classified as spam. This Bayesian approach to spam is highly effective, a May 2003 BBC article reported that spam detection rates of over 99.7 percent can be achieved with a very low number of false positives.

Conclusion
Bayesian filtering, if implemented the right way and tailored to an organization, is the most effective technology to combat spam. The downside to this technique is that, we have to wait at least two weeks for it to learn and create the ham or spam database. Nevertheless, over time, the Bayesian filter becomes more and more effective as it learns about the organization’s email habits, along with updating through other anti-spam databases. 

Source:
White Paper, GFI (2008). Why Bayesian filtering is the most effective anti-spam technology. Retrieved from http://www.gfi.com/whitepapers/why-bayesian-filtering.pdf

Monday, April 16, 2012

Measuring Sustained Competitive Advantage Using Bayesian Reasoning

Introduction:
In this article, Tang and Liou develop a theoretical framework to understand the causal relationships among (1) sustainable competitive advantage, (2) configuration, (3) dynamic capability, and (4) sustainable superior performance. They propose that a firm’s competitive advantage, resource bundle configuration, and dynamic learning capability cannot be comprehended by outsiders. Its operational performance, however, can be captured by financial indicators.

They promote an inductive Bayesian interpretation of the sustainable competitive advantage proposition. From this viewpoint, the presence or absence of competitive advantage may be reflected in the causal relationship between resource configuration, dynamic capability, and observable financial performance. They then apply this theoretical framework to an example drawn from the global semiconductor industry, an area in which resource configuration and dynamic capability are essential to performance.

Summary:
The paper expounds on the theory supporting the use of Bayesian reasoning to measure competitive advantage over other measures like Porter’s competitive strategy or the resource based view of valuable, rare, inimitable and non-substitutable.

Powell’s premise – sustainable competitive advantage is more probable in firms that have already achieved sustained superior performance – is further developed through periodically updating its propositions or hypotheses in the face of empirical evidence. Tang and Liou use a Bayesian discriminate model to reveal the functional dependence of superior performance on unique business processes. The primary sources of competitive advantage are considered embedded in and inseparable from the organization itself, along with its business units and functional departments. It is assumed that the process of managing these resources, termed strategic fit, cannot be comprehended or imitated by outsiders. The model is then demonstrated using the semi-conductor industry.
Explanation of sustainable competitive advantage
Model:
The final data set for the study contained 147 companies and 786 firm-year observations. Of those, 188 companies are located in developed countries (The U.S., within Europe, and Japan). The other 29 are in the Asia/Pacific region. Using the firm’s financial data, certain observable traits can be inferred.

The study began by using PCA, principle component analysis, on the financial indicators to identify the traits or factors. Three principal factors accounted for 60 percent of the total variance.

Factor 1: Relationship management. This factor includes customer relationship management (accounts receivable turnover), three variables related to supplier relationship management (accounts payable turnover, inventory turnover, and CGS/sales) and one variable associated with the government (tax to sales ratio). The factor illustrates the sustainable competitive advantage of firms that manage upstream (suppliers), downstream (customers), and governmental relationships. The variance indicated that good relationship management can pay off with respect to a lower CGS.

Factor 2: Management ability. This factor consists of indicators related to fixed asset management capabalities including department/sales ratio and fixed asset turnover. The correlation between fixed assets turnover and Factor 2 indicates that firms with greater competence in assets management generate revenue at a lower unit cost and low asset depreciation.

Factor 3: Knowledge management. This factor includes R&D/sales and SG&A/sales ratios to measure a firm’s effectiveness in resource deployment. The high correlation indicates that lower unit costs are associated with efficient management.

The findings, therefore, support the idea that resource configurations or factors of a firm can be inferred from their observable financial indicators.

Conclusion:
Tang and Liou advance Powell’s idea of using Bayesain probabilistic reasoning as a means of distinguishing sustained competitive advantage from sustained superior performance in this paper.
They propose that particular resource configurations can be shown to link the two – sustained competitive advantage and sustained superior performance.

Through a discussion of Bayes’ theory and subsequent semi-conductor example, the paper describes how empirical data on past financial performance in a population of firms can be used to generate the posterior probability of sustainable competitive advantage, given the prior probabilities of both competitive advantage and competitive disadvantage.

Source:Tan, Y-C; Liou, F-M. (2010). Does Firm Performance Reveal Its Own Causes? The Role of Bayesian Inference. Strategic Management Journal. 31: 39-57.
Retrieved from http://web.it.nctu.edu.tw/~etang/SMJ2010_TangLiou.pdf

Saturday, April 14, 2012

A Test of Empirical Bayes Journey-to-Crime Estimation in The Hague


Introduction:
In the article, “Finding a Serial Burglar’s Home Using Distance Decay and Conditional Origin–Destination Patterns: A Test of Empirical Bayes Journey-to-Crime Estimation in The Hague”, the authors test a new method, empirical Bayes journey-to-crime estimation, to estimate where an offender lives from where he or she commits crimes. In the new method, the profiler not only asks ‘what distances did previous offenders travel between their home and the crime scenes’ but also ‘where did previous offenders live who offended at the locations included in the crime series I investigate right now?’.

Summary:
The empirical Bayes method uses not only the distance of the journey to crime, but also exploits our knowledge of origins (where did previous offenders live) and destinations (where did they offend), and the links between them to predict the home of a serial offender. It uses more specific information about past offenders. In contradistinction from previous methods, distance does not completely dictate the outcome of the prediction. Thus, given distance, if some destinations have been associated with a particular origin relatively frequently in the past, the new method will identify that particular origin as a likely home area of the offender.

The Bayes journey-to-crime estimation is an extension of its earlier distance-based. Based on connections between offenders and the incidents they committed, three risk surfaces are calculated:
  1. The first risk surface is the risk surface generated by the regular journey-to-crime/distance decay      method in CrimeStat (labeled distance decay risk surface & P(JTC)).
  2. The second is a ‘usual suspects’ risk surface by prioritizes zones where previous offenders lived, independent of where they committed their crimes and independent of the location of the crimes in the series of the offender that is being searched (labeled the general risk surface & P(O))
  3. The third risk surface is based on the origin-destination zone matrix (labeled conditional probability risk surface & P(O|JTC)).
The empirical Bayes journey-to-crime method generates two other risk surfaces by combining the above three risk surfaces. One of these two combination surfaces is the product risk surface, which explicitly recognizes both distance decay and the home-to-incident histories of prior offenders. The product surface is mathematically the numerator of the other combination surface, the Bayesian risk probability. The Bayesian risk surface is calculated by the application of the Bayes’ formula:
Thus, in addition to the three basic risk surfaces distance decay, general, and conditional, in this paper, two combination risk surfaces, product and Bayesian risk, are analyzed.


Conclusion:
Based on the study of 62 burglars, the homes of serial burglars were more successfully estimated with the new conditional risk surface than with the other risk surfaces. The method demonstrated may seem complex, but the authors are confident that in practice it will not be. The authors do state that there are disadvantages with using Bayes method which includes that the new method requires more data, more upfront work and may only be applied to relatively common crimes such as burglary or robbery.

Source:
Block, R., & Bernasco, W. (2009). Finding a serial burglar's home using distance decay and conditional origin–destination patterns: a test of empirical Bayes journey-to-crime estimation in the Hague. Journal Of Investigative Psychology & Offender Profiling, 6(3), 187-211. doi:10.1002/jip.108. Received from http://web.ebscohost.com/ehost/pdfviewer/pdfviewer?sid=5720f7fe-4ccf-473d-b299-e8755df7f04b%40sessionmgr113&vid=6&hid=110

Monday, May 4, 2009

Bayesian statistics: principles and benefits

http://library.wur.nl/frontis/bayes/03_o_hagan.pdf

This article is meant to summarize the basics of Bayesian statistics for beginners.

In Bayesian statistics:
Graphically, the narrower the curve, the tighter the parameters. The difference between frequentist methods and Bayesian analysis is the use of past information, which is principally subjective. It is important for the prior information to be defensible and reasonable. The author believes subjectivity is a strength of the system because it allows for the examination of posterior distributions from different informed observers.

Until the 1990s, computational tools for conducting Bayesian analysis were nascent or non-existent. While there are tools for the specialist available today, the general practitioner of Bayesian analysis will find there are few user friendly tools available.

The author enumerates a number of benefits for using Bayesian analysis. They are:
  • It provides meaningful and intuitive inferences.
  • It can answer complex questions cleanly and exactly.
  • It makes use of all available information.
  • It is well suited for decision-making.
Enhanced by Zemanta