Search engines provide an interface to a group of items that enables users to specify criteria about an item of interest and have the engine find the matching items. The criteria are referred to as a search query. In the case of text search engines, the search query is typically expressed as a set of words that identify the desired concept that one or more documents may contain.[1]
There are several styles of search query syntax that vary in strictness. Where as some text search engines require users to enter two or three words separated by white space, other search engines may enable users to specify entire documents, pictures, sounds, and various forms of natural language. Some search engines apply improvements to search queries to increase the likelihood of providing a quality set of items through a process known as query expansion.
index-based search engineThe list of items that meet the criteria specified by the query is typically sorted, or ranked, in some regard so as to place the most relevant items first. Ranking items by relevance (from highest to lowest) reduces the time required to find the desired information. Probabilistic search engines rank items based on measures of similarity and sometimes popularity or authority. Boolean search engines typically only return items which match exactly without regard to order.
To provide a set of matching items quickly, a search engine will typically collect metadata about the group of items under consideration beforehand through a process referred to as indexing. The index typically requires a smaller amount of computer storage, and provides a basis for the search engine to calculate item relevance. The search engine may store of copy of each item in a cache so that users can see the state of the item at the time it was indexed or for archive purposes or to make repetitive processes work more efficiently and quickly.
Notably, some search engines do not store an index. Crawler, or spider type search engines may collect and assess items at the time of the search query. Meta search engines simply reuse the index or results of one or more other search engines.
source: http://en.wikipedia.org/wiki/Search_engines
Showing posts with label Theory. Show all posts
Showing posts with label Theory. Show all posts
Tuesday, November 20, 2007
Monday, October 29, 2007
Theory
Statistical analysis of website usage
Analysing data from your website is important. Perception is often vastly different to reality. Website statistics can be misleading if not interpreted properly. A basic analysis can be done using statistics programs provided by your hosting company. Using these statistics, you should be able to evaluate;
Return on investment in SEO. Search engine optimisation companies who charge thousands of dollars for their services often base their claim on getting a handful of top search terms but it may be that only a handful of visitors actually find your website by typing in those search terms. Are you spending hundreds of dollars on hosting, management, search engine optimisation and copywriting a month for a website that no-one visits?
Are your premium google listings, google adwords, adverting costs from other search engines and websites providing a return on investment, or are you spending thousands of dollars for a handful of visitors?
Where are your visitors coming from?
What are the popular search terms and phrases used to find your website?
Are marketing campaigns like letter box drops, competitions, raffles, etc bringing more visitors to your website?
Limitations
Website data like all data collection has limitations. Currently all website statistical programs and web services on the internet assume that a unique user is equal to a unique I.P. Users may choose to access your website from different locations and/or may not have a static IP. Users of dial up internet will not have a static IP and so the website data will count them as a different visitor. Understanding the limitations, is important when interpreting data.
Many people believe that a popular website gets lots of hits. The number of hits is the equal to the number of file requests on a webpage, so if you have lots of images on your website, you will have lots of hits. The number of hits tells you nothing about the popularity of your website or whether your website is working for your business.
All the log analyzers and statistics programs available use simple mathematics. The problem is averages can lie if the proper mathematical model is not used. A statistician knows that averages, medians are meaningless without discussing variance, standard deviations, testing hypothesis, choosing the correct distribution, removing outliners.
Finding meaning in statistics
Statistics is not flawless but it will paint a picture and a good analysis of website usage will benchmark website performance against your desired outcomes and provide options for better website design.
Some questions statistical analysis may provide answers to include -
How many visitors are coming to the website?
Where do visitors come from? If they come from search engines, what key words and phrases are they using to find your website?
How long do visitors stay on your website?
What are the click paths of various type of visitors to your website?
What linking sites do your visitors come from?
Do the statistics give an indication of demographics?
What is the main reason people come to the website to do, read or buy?
However if you want to measure improvement, you need to do a proper analysis and calculate standard deviation. Otherwise, you could prematurely come to a conclusion that there had been an improvement when perhaps it was only variance or seasonal factor that influenced the result.
Using raw logs
Proper statistical analysis can only be done using raw data and by importing data into mathematical programs or log analysing software and eliminating outliners.eg. eliminating the ip addresses of staff members from website usage data. Obviously this is more expensive and the scope and purpose of website statistical analysis must be determined. It is also necessary to compare various tools, as the way that each program manipulates the data will vary and some log analyzers may be programmed incorrectly. Just because it spits out an answer or pie graph does not mean that the answer is correct.
source: http://www.passioncomputing.com.au/Website_Redesign/Statistical_analysis.aspx
Analysing data from your website is important. Perception is often vastly different to reality. Website statistics can be misleading if not interpreted properly. A basic analysis can be done using statistics programs provided by your hosting company. Using these statistics, you should be able to evaluate;
Return on investment in SEO. Search engine optimisation companies who charge thousands of dollars for their services often base their claim on getting a handful of top search terms but it may be that only a handful of visitors actually find your website by typing in those search terms. Are you spending hundreds of dollars on hosting, management, search engine optimisation and copywriting a month for a website that no-one visits?
Are your premium google listings, google adwords, adverting costs from other search engines and websites providing a return on investment, or are you spending thousands of dollars for a handful of visitors?
Where are your visitors coming from?
What are the popular search terms and phrases used to find your website?
Are marketing campaigns like letter box drops, competitions, raffles, etc bringing more visitors to your website?
Limitations
Website data like all data collection has limitations. Currently all website statistical programs and web services on the internet assume that a unique user is equal to a unique I.P. Users may choose to access your website from different locations and/or may not have a static IP. Users of dial up internet will not have a static IP and so the website data will count them as a different visitor. Understanding the limitations, is important when interpreting data.
Many people believe that a popular website gets lots of hits. The number of hits is the equal to the number of file requests on a webpage, so if you have lots of images on your website, you will have lots of hits. The number of hits tells you nothing about the popularity of your website or whether your website is working for your business.
All the log analyzers and statistics programs available use simple mathematics. The problem is averages can lie if the proper mathematical model is not used. A statistician knows that averages, medians are meaningless without discussing variance, standard deviations, testing hypothesis, choosing the correct distribution, removing outliners.
Finding meaning in statistics
Statistics is not flawless but it will paint a picture and a good analysis of website usage will benchmark website performance against your desired outcomes and provide options for better website design.
Some questions statistical analysis may provide answers to include -
How many visitors are coming to the website?
Where do visitors come from? If they come from search engines, what key words and phrases are they using to find your website?
How long do visitors stay on your website?
What are the click paths of various type of visitors to your website?
What linking sites do your visitors come from?
Do the statistics give an indication of demographics?
What is the main reason people come to the website to do, read or buy?
However if you want to measure improvement, you need to do a proper analysis and calculate standard deviation. Otherwise, you could prematurely come to a conclusion that there had been an improvement when perhaps it was only variance or seasonal factor that influenced the result.
Using raw logs
Proper statistical analysis can only be done using raw data and by importing data into mathematical programs or log analysing software and eliminating outliners.eg. eliminating the ip addresses of staff members from website usage data. Obviously this is more expensive and the scope and purpose of website statistical analysis must be determined. It is also necessary to compare various tools, as the way that each program manipulates the data will vary and some log analyzers may be programmed incorrectly. Just because it spits out an answer or pie graph does not mean that the answer is correct.
source: http://www.passioncomputing.com.au/Website_Redesign/Statistical_analysis.aspx
Labels:
IUOMA,
Ruud Janssen,
Statictical Data,
Theory
Wednesday, September 19, 2007
Technical Details eXTReMe Tracker
The Tracker-code has to be copied into the source of your page integrally and without *any* changes. Only then the tracking will be accurate. If you change the code eXTReMe digital shall have the right to suspend services. The button included within the code must be visible on your page and in the size as exposed.
With the eXTReMe Tracker you get every advanced feature required to picture the visitors of your website. Conveniently arranged, numbers, percentages, stats, totals and averages. All the way up from simple counting your visitors until tracking the keywords they use to find you.
Information you can't website without!
service:
• Real-time reporting
• Wide range of specifications
• Extensive referrer tracking
• JavaScript optimized (5 times faster)
• Your time zone
• No traffic limitations
• Setup within a few minutes
• Completely FREE!
Featuring:
• Summary
Totals and Averages
• Basic Tracking 1
Unique Visitors:
- Days
- Weeks
- Months
- Hours of the day
- Days of the Week
• Basic Tracking 2
Incl., Excl. Reloads:
- Days
- Weeks
- Months
• Geo Tracking
Domains
Countries
Continents
• System Tracking
Browsers
JavaScript Enabled
Operating Systems
Screen Resolutions
Screen Colors
• Referrer Tracking 1
Last 20
Last 20 from Email
Last 20 from Searchengines
Last 20 Queries
Last 20 from Usenet
Last 20 from Harddisk
• Referrer Tracking 2
Totals by Source:
- Website
- Searchengine \
- Email
- Usenet
- Harddisk
Totals by Searchengine:
- 24 most popular engines
All Keywords
All Website Referrers
With the eXTReMe Tracker you get every advanced feature required to picture the visitors of your website. Conveniently arranged, numbers, percentages, stats, totals and averages. All the way up from simple counting your visitors until tracking the keywords they use to find you.
Information you can't website without!
service:
• Real-time reporting
• Wide range of specifications
• Extensive referrer tracking
• JavaScript optimized (5 times faster)
• Your time zone
• No traffic limitations
• Setup within a few minutes
• Completely FREE!
Featuring:
• Summary
Totals and Averages
• Basic Tracking 1
Unique Visitors:
- Days
- Weeks
- Months
- Hours of the day
- Days of the Week
• Basic Tracking 2
Incl., Excl. Reloads:
- Days
- Weeks
- Months
• Geo Tracking
Domains
Countries
Continents
• System Tracking
Browsers
JavaScript Enabled
Operating Systems
Screen Resolutions
Screen Colors
• Referrer Tracking 1
Last 20
Last 20 from Email
Last 20 from Searchengines
Last 20 Queries
Last 20 from Usenet
Last 20 from Harddisk
• Referrer Tracking 2
Totals by Source:
- Website
- Searchengine \
- Usenet
- Harddisk
Totals by Searchengine:
- 24 most popular engines
All Keywords
All Website Referrers
Labels:
Extreme Tracker,
Ruud Janssen,
Statistics,
Theory
PageRanks - Theory
Google searches more sites more quickly, delivering the most relevant results.
Introduction
Google runs on a unique combination of advanced hardware and software. The speed you experience can be attributed in part to the efficiency of our search algorithm and partly to the thousands of low cost PC's we've networked together to create a superfast search engine.
The heart of our software is PageRank™, a system for ranking web pages developed by our founders Larry Page and Sergey Brin at Stanford University. And while we have dozens of engineers working to improve every aspect of Google on a daily basis, PageRank continues to play a central role in many of our web search tools.
PageRank Explained
PageRank relies on the uniquely democratic nature of the web by using its vast link structure as an indicator of an individual page's value. In essence, Google interprets a link from page A to page B as a vote, by page A, for page B. But, Google looks at considerably more than the sheer volume of votes, or links a page receives; for example, it also analyzes the page that casts the vote. Votes cast by pages that are themselves "important" weigh more heavily and help to make other pages "important." Using these and other factors, Google provides its views on pages' relative importance.
Of course, important pages mean nothing to you if they don't match your query. So, Google combines PageRank with sophisticated text-matching techniques to find pages that are both important and relevant to your search. Google goes far beyond the number of times a term appears on a page and examines dozens of aspects of the page's content (and the content of the pages linking to it) to determine if it's a good match for your query.
Integrity
Google's complex automated methods make human tampering with our search results extremely difficult. And though we may run relevant ads above and next to our results, Google does not sell placement within the results themselves (i.e., no one can buy a particular or higher placement). A Google search provides an easy and effective way to find high-quality websites that contain information relevant to your search.
©2007 Google , source : http://www.google.nl/technology/
Introduction
Google runs on a unique combination of advanced hardware and software. The speed you experience can be attributed in part to the efficiency of our search algorithm and partly to the thousands of low cost PC's we've networked together to create a superfast search engine.
The heart of our software is PageRank™, a system for ranking web pages developed by our founders Larry Page and Sergey Brin at Stanford University. And while we have dozens of engineers working to improve every aspect of Google on a daily basis, PageRank continues to play a central role in many of our web search tools.
PageRank Explained
PageRank relies on the uniquely democratic nature of the web by using its vast link structure as an indicator of an individual page's value. In essence, Google interprets a link from page A to page B as a vote, by page A, for page B. But, Google looks at considerably more than the sheer volume of votes, or links a page receives; for example, it also analyzes the page that casts the vote. Votes cast by pages that are themselves "important" weigh more heavily and help to make other pages "important." Using these and other factors, Google provides its views on pages' relative importance.
Of course, important pages mean nothing to you if they don't match your query. So, Google combines PageRank with sophisticated text-matching techniques to find pages that are both important and relevant to your search. Google goes far beyond the number of times a term appears on a page and examines dozens of aspects of the page's content (and the content of the pages linking to it) to determine if it's a good match for your query.
Integrity
Google's complex automated methods make human tampering with our search results extremely difficult. And though we may run relevant ads above and next to our results, Google does not sell placement within the results themselves (i.e., no one can buy a particular or higher placement). A Google search provides an easy and effective way to find high-quality websites that contain information relevant to your search.
©2007 Google , source : http://www.google.nl/technology/
Labels:
Google,
PageRank,
Ruud Janssen,
Statistics,
Theory
Subscribe to:
Posts (Atom)
