
Sunday, November 16, 2008
Stats on 16-11-2008

Saturday, January 19, 2008
Wednesday, December 12, 2007
Data Mining - Theory
`We are drowning in information but starved for knowledge.'
John Naisbitt
What is data mining?
Data mining sits at the interface between statistics, computer science, artificial intelligence, pattern recognition, machine learning, database management and data visualisation (to name some of the fields).
Data mining is the non-trivial process of identifying valid, novel, potentially useful, and ultimately comprehensible patterns or models in data to make crucial decisions. Data mining is not a product that can be bought. Data mining is a discipline and process that must be mastered - a whole problem solving cycle.
The main part of data mining is concerned with the analysis of data and the use of software techniques for finding patterns and regularities in sets of data. The idea is that it is possible to strike gold in unexpected places as the data mining software extracts patterns not previously discernible or so obvious that no-one has noticed them before. The analysis process starts with a set of data, uses a methodology to develop an optimal representation of the structure of the data during which time knowledge is acquired. Once knowledge has been acquired this can be extended to larger sets of data working on the assumption that the larger data set has a structure similar to the sample data. This is analogous to a mining operation where large amounts of low grade materials are sifted through in order to find something of value.
Is data mining `statistical déjà vu'?
Whereas statistical analysis traditionally concerns itself with analysing primary data that has been collected to check specific research hypotheses (`primary data analysis'), data mining can also concern itself with secondary data collected for other reasons (`secondary data analysis'). Furthermore, data can be experimental data (perhaps the result of an experiment which randomly allocates all the statistical units to different kinds of treatment), but in data mining the data is typically observational data.
Data warehousing provides the enterprise with a memory
Companies are collecting data on seemingly everything. For example, a customer-focused enterprise regards every record of an interaction with a client or prospect (e.g. each call to customer support, each point-of-sale transaction, each catalogue order, each visit to a company web site) as a learning opportunity. But, learning requires more than simply gathering data. In fact, many companies gather hundreds of gigabytes of data without learning anything. For example, data are gathered because they are needed for some operational purpose, such as inventory control or billing. Once data served that purpose, data languish on tape or get discarded. The data's hidden value has largely gone untapped. For learning to take place, data from many sources (e.g. billing records, scanner data, registration forms, applications, call records, coupon redemption, surveys, manufacturing data) must first be gathered together and organised in a consistent and useful way - in a way that facilitates the retrieval of information for analytic purposes. This is called data warehousing. Data warehousing allows the enterprise to remember what it has noticed about its customers. Data warehousing provides the enterprise with a memory.
Data mining provides the enterprise with intelligence
Memory is of little use without intelligence. That is where data mining comes in. Intelligence allows us to comb through our memories noticing patterns, devising rules, coming up with new ideas to try, and making predictions about the future. The data must be analysed, understood and turned into actionable information. Using several data mining tools and techniques that add intelligence to the data warehouse, you will be able to exploit the vast mountains of data, for example, generated by interactions with your customers and prospects in order to get to know them better. Typical customer-focused business questions are:
What customers are most likely to respond to a mailing?
Are there groups (or segments) of customers with similar characteristics or behavior?
Are there interesting relationships between customer characteristics?
Who is likely to remain a loyal customer and who is likely to jump ship?
Where should the next branch be located?
What is the next product or service this customer will want?
Answers to questions like these lie buried in your corporate data, but it takes powerful data mining tools to get at them, i.e. to dig user info for gold. Data mining provides the enterprise with intelligence. Companies can use data mining findings for more profitable, proactive decision making and competitive advantage.
With data mining, companies can, for example, analyze customers' past behaviors in order to make strategic decisions for the future. Keep in mind, however, that the data mining techniques and tools are equally applicable in fields ranging from law enforcement to radio astronomy, medicine, and industrial process control.
Please contact us today in order to discuss how data mining can be applied to your field or work. Get Statooed.
Data mining myths versus realities
A great deal of what is said about data mining is incomplete, exaggerated, or wrong. Data mining has taken the business world by storm, but as with many new technologies, there seems to be a direct relationship between its potential benefits and the quantity of (often) contradictory claims, or myths, about its capabilities and weaknesses. When you undertake a data mining project, avoid a cycle of unrealistic expectations followed by disappointment. Understand the facts instead and your data mining efforts will be (hopefully) successful. A list of the most common data mining myths versus realities you will find here.
Data mining can not be ignored - the data is there, the methods are numerous, and the advantages that knowledge discovery brings are tremendous. Companies whose data mining efforts are guided by `mythology' will find themselves at a serious competitive disadvantage to those organizations taking a measured, rational approach based on facts.
source: http://www.statoo.com/en/datamining/
Tuesday, December 11, 2007
ClustrMaps - December 11th 2007

Wednesday, October 31, 2007
Motigo - Overview visitors 31-10-2007

While Statcounter only registers the last 500 hits, the Motigo counter saves all hits and therefore becomes interesting when one is searchibf for historical information. The longer a site is in the air, the more interesting the statistical data gets because one can make prognoses. Here is an overview of the countries visitors come from. Seems like they come also from abroad now. Even some dozens of hits from www.google.com themselves (in California, USA), so they are reading this too.....
FEEDJIT - Trafic Map

FEEDJIT - Live Traffic Feed
FeeJit has a nice result besides the wesite that shows how visitors came here and where to they leave. The overview is visible for everybody so they can see what kind of 'traffic' this site genereates. This is also for the owner of a site very interesting information because when he checks the site he sees what is going on.About FEEDJIT
"Our mission is to provide high performance real-time widgets for the blogging community that are free and easy to use. FEEDJIT is founded by two serial entrepreneurs whose previous businesses include a search engine and a blogging platform. If you'd like to contact us, please email"
<support@feedjit.com> with any questions, feedback or bug reports.
Monday, October 29, 2007
FEEDJIT
The above textcomes from their own site.
Details on: http://feedjit.com/
I am testing this new application now on this statistical blog as well. It is a Widget that shows the visitors how others get on your site. Other statistical programms have this information as well but mostly it isn't accesible for the visitors. Visitors can follow the same links as previous visitors.
Theory
Analysing data from your website is important. Perception is often vastly different to reality. Website statistics can be misleading if not interpreted properly. A basic analysis can be done using statistics programs provided by your hosting company. Using these statistics, you should be able to evaluate;
Return on investment in SEO. Search engine optimisation companies who charge thousands of dollars for their services often base their claim on getting a handful of top search terms but it may be that only a handful of visitors actually find your website by typing in those search terms. Are you spending hundreds of dollars on hosting, management, search engine optimisation and copywriting a month for a website that no-one visits?
Are your premium google listings, google adwords, adverting costs from other search engines and websites providing a return on investment, or are you spending thousands of dollars for a handful of visitors?
Where are your visitors coming from?
What are the popular search terms and phrases used to find your website?
Are marketing campaigns like letter box drops, competitions, raffles, etc bringing more visitors to your website?
Limitations
Website data like all data collection has limitations. Currently all website statistical programs and web services on the internet assume that a unique user is equal to a unique I.P. Users may choose to access your website from different locations and/or may not have a static IP. Users of dial up internet will not have a static IP and so the website data will count them as a different visitor. Understanding the limitations, is important when interpreting data.
Many people believe that a popular website gets lots of hits. The number of hits is the equal to the number of file requests on a webpage, so if you have lots of images on your website, you will have lots of hits. The number of hits tells you nothing about the popularity of your website or whether your website is working for your business.
All the log analyzers and statistics programs available use simple mathematics. The problem is averages can lie if the proper mathematical model is not used. A statistician knows that averages, medians are meaningless without discussing variance, standard deviations, testing hypothesis, choosing the correct distribution, removing outliners.
Finding meaning in statistics
Statistics is not flawless but it will paint a picture and a good analysis of website usage will benchmark website performance against your desired outcomes and provide options for better website design.
Some questions statistical analysis may provide answers to include -
How many visitors are coming to the website?
Where do visitors come from? If they come from search engines, what key words and phrases are they using to find your website?
How long do visitors stay on your website?
What are the click paths of various type of visitors to your website?
What linking sites do your visitors come from?
Do the statistics give an indication of demographics?
What is the main reason people come to the website to do, read or buy?
However if you want to measure improvement, you need to do a proper analysis and calculate standard deviation. Otherwise, you could prematurely come to a conclusion that there had been an improvement when perhaps it was only variance or seasonal factor that influenced the result.
Using raw logs
Proper statistical analysis can only be done using raw data and by importing data into mathematical programs or log analysing software and eliminating outliners.eg. eliminating the ip addresses of staff members from website usage data. Obviously this is more expensive and the scope and purpose of website statistical analysis must be determined. It is also necessary to compare various tools, as the way that each program manipulates the data will vary and some log analyzers may be programmed incorrectly. Just because it spits out an answer or pie graph does not mean that the answer is correct.
source: http://www.passioncomputing.com.au/Website_Redesign/Statistical_analysis.aspx
Friday, October 26, 2007
GeoVisite

Thursday, October 25, 2007
StatCounter - Returning Visits

This overview is always interesting. Do visitors come back after a first visit. Seems that is the case here, judging these graphical presented results by www.statcounter.com
StatCounter - Returning Visits
From your project log of the last x number of pageloads, we extract the total number of unique visitors present in it. Each unique visitor has a cookie, which is incremented each time they return to visit your website (a couple of hours is needed between visits depending on your settings). From this info we can show you how often visitors return to see your website again and again.
The best and most successful websites are the ones with a very high return frequency. If you have a low or non-existant return frequency you may want to change your website to encourage your visitors to come again and again.
Statcounter - Popular Pages
This overview of www.statcounter.com gives you a perfect overview of what the popular pages on your website are. These subjects are interesting to your visitors!
Statcounter - Popular Pages
You will quickly see which pages are the most heavily visited by your visitors and what ones are being left well enough alone. If some pages are being overlooked it could be a good idea to improve or make the navigation to those pages more obvious. Or entice your visitors to find out more about these pages.
Drill down the data - show all your visitors that visited this particular page during their visit!
Common problem with Popular Page Stats
If you only install the StatCounter code on one page of your website we can only track one page of your website. It is highly recommended to install the same code on all pages of your website you want to track.
Queries - Extreme Tracker
Yes, I must admit I am a fan of this Extreme Tracker because it produces overviews that are very interesting. This is just a sample: with which words do the visitors come to my site and through which searchengine with what keywords? I like the answers to these questions. Seems like more people do since they ask these questions.Friday, October 19, 2007
ClustrMap Results on 19-10-2007

Running total of visits to the above URL since 1 Oct 2007: 141Total since archive, i.e. 1 Oct 2007 - present: 141 (not necessarily all displayed - see below). Visits on previous 'day': 1.Additional Notes about totals and map updates (for full guide see Map Key):The map shows individual visits to the web site shown at the top of the page, clustered within a given distance. The location of each visit is based on the IP address of the computer used. Update frequency variation: In order for your map to be 'updated' (whether daily, weekly, or monthly) the number of visitors shown must also have grown by a certain percentage since the last update.
This percentage may have changed recently, and is explained more fully in our FAQ about update frequency.Total/subtotal discrepancies above are typically caused either by IP addresses not currently being in the database or by the gap between counter tallies (continous) and map updates (which may be daily, weekly, or monthly).
See FAQ for additional explanations.
Friday, October 12, 2007
Statistics Cartoon

Tuesday, October 9, 2007
Saturday, October 6, 2007
ClustrMap on 6-10-2007

Keywords - StatCounter: List

StatCounter gives you a list of keywords people looked for and then found your site. Here you can see how people came to this actual Statistical Data site. It shows the interest of you visitors.


