Showing posts with label heteroscedasticity. Show all posts
Showing posts with label heteroscedasticity. Show all posts

Monday, 22 January 2018

Data Handling: Interpretation and Discussion of Results in Scientific Economic Research

Philip O. Alege, Ph.D
Professor of Economics
                                 Department of Economics and Development Studies                                
Covenant University, Ota, Ogun State

Introduction
The tools available to modern economists in the discharge of functions as an analyst are very many simply because of various infiltrations of knowledge from other sciences into the discipline of economics such as physics, biology, mechanical engineering and particularly mathematics and statistics. Today, modern economies will be difficult to analyse, understand and predict without the tools of mathematics, statistics and in particular econometrics. This can be explained by virtue of the growing number of economic activities and interactions among the different agents in a given country and between/among countries. There is a school of thought that believes in more of economics and little of mathematics. There is also another school of thought that believes in substantial application of the tools of mathematics in economics as necessary to get the “useful” results from our analysis. Though I belong to the latter school, I do also contend that things must be done properly.

Basics of Econometric Modeling
Basically, econometrics is to provide empirical support for economic data. Its main purpose is to estimate the parameter(s) of a model that capture the behaviour of economic agent(s) as described by the theory and the model. Since the estimated parameters may be useful in understanding the economic theory, for policy analysis and forecasting, it becomes necessary on the econometrician to obtain parameters that are efficient. In order to achieve this, we must adhere to some principles of model building that can generate results whose interpretations and discussions will be useful for policy analysis as well as decision making. These are listed as follows:
·    Economic theory applicable to the specific area of the research
·    Design of the mathematical model and the hypotheses of the study
· The quest to obtain the right economic statistics i.e. the collection, collation and analysis of requisite data for the research, and
·  Interpretation/discussion of the findings/results

The researcher should keep in mind that econometric models are tools and therefore means to some desired ends. That is, our professional calling is to provide plausible parameter estimates that should be useful for policy analysis and decision making. Therefore, any mathematical and/or statistical model must be able to deliver these objectives of the researcher in an efficient manner.

Model Specification and Estimation Techniques
Consequently, model specification is the nucleus/DNA of any scientific economic research. I usually call it the economics of the study. It shows the depth of the researcher in the knowledge of theoretical economics as well as ability to state clearly the contribution(s) to knowledge as envisaged in the study. The latter may come as:
·      Additional variable to existing theoretical model
·  A single equation now specified as system of equations in order to capture a phenomenon hitherto not considered, or
·      Application of a technique not commonly used in our own environment.

Once the model is correctly specified, the next step is to consider the estimation technique that will produce the most efficient estimates of the parameters of the model. It is important to use a technique of estimation that will deliver the objective(s) of the study. It is apposite to mention some estimation techniques at this stage. It should, however, be noted that the list is not exhaustive. Some of these are as follows: ordinary least squares (OLS), indirect least squares (ILS), instrumental variables (IL), two stage least squares (2SLS): in the case of system of simultaneous equation, three stage least squares (3SLS), error correction model (ECM) which examines short-run dynamics, cointegration regression, generalised method of moments (GMM), vector autoregressive method (VAR) which examines the effect of shocks on a system, structural vector autoregressive method (SVAR), panel data method, panel vector autoregressive method (PVAR), panel structural vector autoregressive method (PSVAR), vector error correction (VECM), panel cointegration, panel vector error correction (PVECM) and so on.
Some learning resource materials are, but not limited to:
1.    Gujarati D. N. (2013). Basic Econometrics, Eight Edition, McGraw-Hill International Editions Economic Series, Glasgow
2.    Maddala, G. S. and Lahiri, K. (2009). Introduction to Econometrics. Fourth Edition. John Wiley
3.    Wooldridge J. M. (2009). Introductory Econometrics, Fourth Edition, South-Western Cengage Learning, Mason, U.S.A

Dynamic General Equilibrium (DGE) Models
There are lots of other techniques of estimation that should be of interest to the younger generations of economists. The basic framework is the dynamic general equilibrium (DGE) theories. Models built around this method are solved using the DYNARE codes in the MATLAB environment or directly using the Matlab codes written for such models. As part of the estimation is the need to calibrate the model. This consists of finding values for some parameters in the model though theoretical knowledge, calculating long-run averages as well as micro-econometric studies. The statistics often used are derived from the Bayesian inference as against the classical statistics referred to in the preceding paragraphs. Some of these models are: real business cycle (RBC), New Keynesian models (NKM), dynamic stochastic general equilibrium (DSGE), over-lapping generation (OLG), computable general equilibrium (CGE), dynamic computable general equilibrium (DCGE), Bayesian vector autoregression (BVAR), Bayesian structural vector autoregression (BSVAR), dynamic macro panels (DMP), augmented gravity models (AGM) and multicounty New Keynesian (MCNK) models.
Some learning resource materials are:
1.    Wichens, M. (2008). Macroeconomic Theory: A Dynamic General Equilibrium Approach. Princeton University Press, Princeton
2.    Canova, F. (undated). Methods for applied Macroeconomic Research
3.    Dejong, D. N. and Dave, C. (2007). Structural Macroeconometrics, Princeton University Press, Princeton.
4.    Cooley, T. F. (ed.) (1995). Frontiers of Business Cycle Research. Princeton University Press, Princeton.
5.    McCandless, G. (2008). The ABCs of RBCs: An Introduction to Dynamic Macroeconomic Models. Harvard University Press; and
6.    Lucas, R. E. (1991). Models of Business Cycles.

It is apposite to state that researchers must have a working understanding of the tests that must be carried out under each technique of estimation. I need to also draw the attention of interested researcher in the area of dynamic general equilibrium because it requires adequate knowledge of computational economics. Specifically, you need sound working knowledge of the following: dynamic optimization, method of Lagrange multipliers, continuous-time optimization, dynamic programming, stochastic dynamic optimization, time-consistency and time-inconsistency and linear rational-expectation models.
Some learning resource materials
1.    Dadkhah, K. (undated) Foundation of Mathematical and Computational Economics, Thomson South-Western.

Interpretation of Results
In interpreting the results of an econometric model, you have the choice of the most appropriate method for your work either the classical or Bayesian statistics as mentioned above. This aspect of the work constitutes the scientific content emanating from economic statistics and mathematical economics. In this case, we should be addressing statistics such as:
·      R-squared
·      Adjusted R-squared (“goodness of fit” test)
·      F-statistics
·      Durbin-Watson statistic
These, in addition to the test of heteroscedasticity constitute the “diagnostic tests”. Once they fail to fall within the zones of acceptance, we cannot go ahead to test for the significance of each variable. There may be the need for: model re-specification, detection and correction of autocorrelation, and/or detection and correction of multicollinearity.

We may also need to test for heteroscedasticity. The occurrence of any of this is an evidence of the violation of assumption(s) of the technique being applied. This is followed by the statistics to test the significance of the individual variables included in the model. This was the standard during the époque of almighty OLS. Later in the history of applied econometrics, it was observed that certain time-series are non-stationary, i.e. their means, variances and covariances are not constant over time. In such situation regression results are generally meaningless and are, therefore, termed spurious. In order to correct for the latter, the statistics often used to examine the stationarity of time series include the following: Dickey-Fuller test, “augmented” Dickey-Fuller test in the presence of error term that is none white noise, Panel data unit root tests, co-integration tests and error correction model (ECM), to mention a few. The use of any of these tests should be in response to the objective of the researcher and the desired contribution(s) to knowledge.

Some Pitfalls in Econometrics
·      The wrong way to go in modelling
How one interprets the coefficients in regression models will be a function of how the dependent (y) and independent (x) variables are measured. In general, there are three main types of variables used in econometrics: (1) continuous variables, (2) the natural logarithm of continuous variables, and (3) dummy variables.



·      Some Specific Rules of Thumb from Statistics

After performing a regression analysis:
1.    Look at the number of observations:
·      Is your result in line with a priori expectation?
·      If not, you should find out why.
·      Remember, any observations with missing values will be dropped from the regression.
·      Do not take the logarithm of a variables whose value equals zero. The model will not run, simple.
·      Ensure the number of observations in your model falls within the rule i.e. sample size should be greater than or equals to 30 (the law of large numbers).

2.    Observe the value of the R2:
·      The R2 tells you the percentage of the total variation in the dependent variable that the independent variables of your model “explains”.
·      This should be less than 1. The rest is the error term.
·      Suppose an estimated model of R2 = 0.46. This means that 46% of the total variation in the dependent variable is explained by the independent variables. This is not a “good fit”.
·      For a regression to have a good fit then we must have a result such that 0.5<R2<1. This is in the case of a time series regression.
·      However, it is considered good for cross-section data and very good for panel data.

Problems with R2:
·      If you have a ‘very low’ R2, have a rethink about whether you might have omitted some important variables.
·      However, be careful not to include unnecessary variables only to increase your R2
·      A ‘very high’ R2 could indicate several problems.
·      Firstly, if a high R2 is combined with many statistically significant variables, your independent variables might be highly correlated amongst themselves (multicollinearity).
·      You might consider dropping some in the interest of parsimony.
·      It might be an indication that you have mis-specified your model.

The adjusted R2:
·      Adjust the R2 to penalize the inclusion of more variables. i.e. correct for the degree of freedom.
·      Include as many variables as you need but keep your model as parsimonious as possible. Observe the rules guiding this.

3.    Look at the F-test.
·      The F-test aims at the “joint significance” of the model.
·      More formally it is a test of whether all your coefficients are jointly equal to zero under the null hypothesis.
·      If they are, effectively your model is not really explaining anything. Hint: ideally you want a high F-value, and a low corresponding p-value 

4.    Interpret the signs of the coefficients.
·      Which ones should be positive and which should be negative from the theoretical perspective? Interpret this!
·      A positive coefficient means that variable has a positive impact on your dependent variable, and a negative one has a negative impact or inverse relationship.

5.    Interpret the size of the coefficients where relevant.
·      If you obtain a statistically significant coefficient-wonderful!
·      So maybe you’ve found consumption increases with disposable income. But by how much? Is it close to 1 by which the marginal propensity to consume is high and the marginal propensity to save is low? What would be the effect of this on the economy?

6.    Look at the significance of the coefficients (most important?).
·      This should in fact become the first thing that your eyes drift towards when you get regression output.
·      You should feel a little hint of excitement as you are waiting to find out whether your model works and whether your theory has been proved correct or not.
·      The test of significance is designed to test whether a coefficient is significantly different from zero or not.
·      If it is not, then you must conclude that your explanatory variable does not, in fact, explain at all your dependent variable.
·      We use t - test (just like we learnt in first year statistics) to test this so that we compare a t - value taken from the table (at a given significance level, α, with n - k degree of freedom) with a calculated t, where n = number of observations and k = number of parameters estimated/independent variables; n – k = degree of freedom.












 


7.    Others
·      Other tests follow, such as testing for normality of error terms, checking for existence of heteroscedasticity, performing specification and robustness tests.
·      But these exciting topics are to be covered if your econometric work would have any useful output valuable for policy making and decision making.

Discussion of Results
The essence of a scientific economic research is to build economic models that enable us obtain plausible estimates from given set of data. We should know that the structural parameters estimated encapsulate our behaviour and, therefore, in discussing them, we need to go beyond the confine of economics to locate additional means of buttressing our results from:
·      historical context
·      socio-political condition, and
·      psychological state as well as
·      international environment

Conclusion
I have tried to raise some important issues in this post. There are so many things to keep in mind when preparing a research work. The most important of them all is the need to keep your model simple and avoid frivolities in modeling. It is important to remember that we are first of all economists. The tools of analysis at our disposal should not overshadow that calling.


Quite a lot has been said about how a researcher in the field of economics can handle data, interpret and discuss research findings that will be relevant for policy-making. If you still have further questions or comments in this regard, kindly post them below for the benefit of all.

Post your comments and questions….

Sunday, 14 January 2018

A Step-by-Step Tutorial on Research and Data Analysis

Note: This tutorial is somewhat detailed!

Data is essential to all disciplines, professions, and fields of endeavours whether in social sciences, arts, technology, life sciences or medicine. The truth is, who are we without data? Data either qualitative or quantitative is informative. It tells us about past and current occurrences. With data, predictions and forecasting can be made either to forestall a negative recurring trend or improve future events. Whichever way, knowing some rules guiding the use of data and how to make it communicate is very important since it often comes out as large voluminous tons of figures or statements. In the same vein, undertaking a research is impossible without data. I mean, what will be the essence of your research if you have no data. In other words, research and data are like siamese twins.

Everyone has different views about how research should be undertaken and how data should be analysed. Afterall, isn’t that why we have different schools of thought? I guess, that’s why. So, what I am about to teach are just simple steps common to all disciplines that are required to undertake any form of research and analyse that data accordingly. Therefore, whether you are a student or a practitioner you will find this guide very helpful. Although, I may be a bit biased towards economics this approach is not fool proof, and regardless of what you know already (and whatever your field is), you will learn a thing or two from this tutorial.

So, let us dig in…..

1.    State the underlying theory.
You must have a theory underlying your study or research. Theories are hypotheses, statements, conjectures, ideas, and assumptions that someone somewhere came up with at some point in time. Such as Darwin’s theory of evolution, Malthusian theory of food and population growth, Keynes’ theory of consumption, McKinnon-Shaw hypothesis on financial reforms etc. Every discipline has its fair share of theories. So make sure you have a theory upon which your research hinges on. It is this theory you are out to test with the available data which culminates into you undertaking a research. Right now, I have a funny theory of my own that countries that have strong and efficient institutions have lower income inequality (…oh well, I just came up with that!). Or yours could be that richer countries have happier citizens. Therefore, anyone can have a theory. Have a theory before you begin that research!

2.    Specify the theoretical (mathematical) model
Having established the theory within which you are situating your research, the next thing to do is to state the theoretical model. Remember, since theories are statements (which are somewhat unobservable), you have to construct them in a functional mathematical form that embodies the theory. The model to be specified is a set of mathematical equations. For instance, given my postulated negative relationship between effective institutions and income inequality, a mathematical economist might specify it as:

                               INQ = b1 + b2INST……………………..[1]

So, equation [1] becomes the mathematical model of the relationship between institutions and income inequality.

Where INQ = income inequality, INST = institutions, b1 and b2 are known as parameters of the model and they are the intercept and slope coefficients. According to the theory, b2 is expected to have a negative sign.

The variable appearing on the left side of the equality sign is the dependent variable or regressand while the one on the right side is called the independent or explanatory variable or regressor. Again, if the model has one equation as it is in equation [1], it is known as a single-equation model and if it has more than one equation, it is called multiple-equation model. Therefore, the inequality-institutions model stated above is a single equation model.

3.    Specify the empirical model
The word “empirical” connotes knowledge derived from experimentation, investigation or verification. Therefore, the mathematical model stated in equation [1] is of limited interest to the econometrician. The econometrician must modify equation [1] to make it suitable for analysis of some sort. This is because, that model assumes that an exact relationship exists between effective institutions and income inequality. However, this relationship is generally inexact. This is because, if we are to obtain institutional data on 10 countries known to have good rankings on governance, rule of law or corruption, we would not expect all their citizens to lie exactly on the straight line. The reason is because aside quality or effective institutions, other variables affect income inequality. Variables such as income level, education, access to loans, economic opportunities etc. are likely to exert some influence on income inequality. Therefore, to capture the inexact relationship(s) between and among economic variables, the econometrician will modify equation [1] as:

                             INQ = b1 + b2INST + u ……………………..[2]

Thus, equation [2] becomes the econometric model of the relationship between institutions and income inequality. It is with this model that the econometrician verifies the inequality-institutions hypothesis using data.

Where u is the disturbance term or often called the error term. The error term is a random variable that may well capture other factors that affect income inequality but not taken into account by the model explicitly. Technically, equation [2] is an example of a linear regression model. The major difference between equations [1] and [2] is that the econometric inequality function hypothesises that the dependent variable INQ is linearly related to the explanatory variable INST but this relationship is not exact due to individual variation represented by u.

4.    Data
Now that you have the theory and have been able to successfully construct your model, the question is, do you have data? To estimate the econometric model stated in equation [2], data is essential to obtain the numerical estimates of b1 and b2. Your choice of data depends on the structure or nature of your research which may determine if you will require the use of qualitative or quantitative data. In line with that, is whether you require the use of primary or secondary data? As a researcher, you can mix both qualitative and quantitative data to gain the breadth and depth of understanding and corroborating what others have done. This is known as meta-data analysis. There is a growing body of researchers using this approach. At this point, you already know whether the data is available for your research or not.

When sourcing your data, identify the dependent variable and the explanatory variables. Let me say a word or two on the explanatory variables. They can further be broken into control variables. The control variables are not directly in your scope of research but they are often included to test if the expected a priori on the key explanatory variable still holds with the inclusion of control variables in the regression model. For instance, using the inequality-institutions model, the dependent variable is INQ, the key explanatory variable is INST and I may decide to control for education, per capita income and access to loans….the last three variables are known as the control variables. Also, in applied research, data is often plagued by approximation errors or incomplete coverage or omitted variables. For instance, social sciences often depend on secondary data and usually have no way of identifying the errors made by the agency that collected the primary data. That being said, do not engage in any research without first knowing that data is available.

...So, start sourcing and putting your data together, we are about to delve into some pretty serious stuff! J

5.    Methodology
The next thing is knowing what methodology to apply. This is peculiar to your research and your discipline. There are so many methodologies, identify the one which best fits your model and use it.

6.    Analytical software
Students often ask me this question: “what analytical software should I use?” My answer has and will always be: “use the software that you are familiar with”. Don’t be lazy! Be proficient in the use of at least one analytical software. There are hundreds of them out there – Stata, R, EViews, SPSS, SAS, Python, SQL, Excel, Agile, and so on. Learn how to use any of them. There are so many tutorial videos on YouTube. For instance, I am very proficient in the use of Stata and Excel analytical softwares with above 60% proficiency in the usage of EViews, SAS and SQL packages. As a researcher and data analyst, you cannot be taken seriously if you cannot lay claim to some level of proficiency in the usage of any of these packages. I use Stata, I love Stata and I will be giving out some periodical hands-on tutorials on how to use Stata to analyse your data. By way of information, I currently analyse data using Stata13.1 package.

So, let us dig in further….it is getting pretty interesting and more involving J

7.    Estimation technique
This is the method of obtaining the estimates for your model, at least an approximation. It is that method based on finding that parameter estimate that best minimises discrepancies between the observed sample(s) and the fitted or predicted model. At this point, you already know what technique to apply that will best give unbiased estimates.

8.    Pre-estimation checks
At this point, you are almost set to begin analysing your data. However, before you proceed, your data must be subjected to some pre-estimation checks. I am very sure that every discipline has these pre-estimation checks in place before carrying out any analysis. In economics there are several of them, such as: multicollinearity test, summary statistics (mean, standard deviation, minimum, maximum, kurtosis, skewness, normality etc.), stationarity test, Hausman test etc. It is from these tests that you identify and correct any abnormality in your data. You may observe the presence of an outlier (when a figure stands out conspicuously either because it is abnormally low or high). You will also get some information regarding the deviation of a variable from the mean (average value), the shape of the probability distribution is also important – is it mesokurtic, platykurtic or leptokurtic? You may want to know whether your data is heavy- or light-tailed. Also, if you are using a time-series data, the stationarity of each variable should be of paramount interest and if it is a panel data (combination of time- and cross-sectional data) the Hausman test should come handy in knowing what estimator (whether fixed or random) to adopt. The bottom-line is that: always carry out some pre-estimation checks before you begin your analysis!

9.    Functional form of the model
Linear relationships are not often common for all economic research, while it is general to come across several studies incorporating many nonlinearities into their regression analysis by simply changing the functional forms of either or both the dependent (regressand) and independent variables (regressors). More often than not, econometricians transform variables from their level forms to functional forms using natural logarithms (denoted as ln). Since variables come in different measurements, it is crucial to know how they are measured in order to make sense of their regression estimates in an equation. For example, using the inequality-institutions model, the inequality variable (using the Gini index) ranges between 0 and 100 and the institution variable is also a decimal ranging between -2.5 and +2.5, obviously these two variables have different measurements. Therefore, an important advantage of transforming variables into natural logarithms (logs, for short) is to equate the variables on the same measurement and applying a constant elasticity relationship and interpretations. It also controls for the presence of outliers in the data amongst others. Let me state here that when the units of measurement of the dependent and independent variables change, the ordinary least squares (OLS) estimates change in entirely expected ways.
Note: changing the units of measurement of only the regressor does not affect the intercept.

Table showing different functional forms of a model
So, given the inequality-institutions model, I may decide to re-specify equation [2] in a log-linear form to obtain an elasticity relationship. That is:

                           lnINQ = b1 + b2lnINST + u ……………………..[3]

10.    Estimate the model
Prior to the existence of analytical softwares, econometricians go through the cumbersome approach of manual computation of regression coefficients. Well, I am glad to tell you that those days are gone forever! With the advent of computerised analytical packages like Stata, EViews, R and the rest of them all you have to do is feed in your data into any that you are familiar with and click “RUN”…and voila! You have your results in split micro-seconds! Most if not all of these packages are excel-friendly. That is, you first have to put your data into an excel format (either .csv, .xls or. xlx file) and then feed into any of them. This is the easiest part of the entire data analysis process. Every researcher loves it whenever they are at this stage. All you need do is feed in your data, click “RUN” and your result is churned out! However, that your coefficients will be according to your expectations is an entirely different story (won’t be told in this write up…hahahaha J).

Below is an example of a result output from Stata analytical software (see heteroscedasticity).


From the regression output, Stata (just like other softwares) provides the beta coefficients, standard errors, t-statistics, probability values, the confidence intervals, R2, F-statistic, number of observations, the degree of freedom, the explained (denoted as Model) and unexplained (denoted as Residual) errors. I will cover analysis of variance (ANOVA) in subsequent tutorials.

11.    Results interpretation
In line with model specification, you can then interpret your results. Always be mindful of the units of measurements (if you are not using a log-linear model). The results output shown above is for a linear (that is, a level-level) model.

12.         Hypothesis testing
Since your primary goal for undertaking the research is to test a theory or hypothesis, it is at this point you do that having stated what your null and alternative hypotheses are. Remember any theory or hypothesis that is not verifiable by empirical evidence is not admissible as a part of scientific enquiry. Now that you have obtained your results, do you reject the null hypothesis in favour of the alternative? The econometric packages always include this element in the result output so that you don’t have to manually compute. Simply check your t-statistics or p-values to know if you will reject the null hypothesis or not.

For instance, from the above output, the beta coefficient for crsgpa is 1.007817 and with the standard error of 0.1038808, you can easily compute the t-statistic as 1.007817/0.1038808 = 9.70 (as given by the Stata output). Importantly, know that a large t-statistic will always provide evidence against the null hypothesis. Likewise, the p-value of 0.000 is indicative of the fact that the likelihood of committing a Type I error (that is, rejecting the null hypothesis when it is true) is very, very remote…close to zero! So, when the null hypothesis is rejected, we say that the coefficient of crsgpa is statistically significant. When we fail to reject the null hypothesis, we say the coefficient is not statistically significant. It is inappropriate to say that you “accept” the null hypothesis. One can only “fail to reject” the null hypothesis. This is because you fail to reject the null hypothesis due to insufficient evidence against it (often due to the sample collected). So, we don’t accept the null, but simply fail to reject it!

(Detailed rudiments of hypothesis testing, Type I and II errors will be covered in subsequent tutorials).

13.  Post-estimation checks
Having obtained your estimates, it is advisable to subject your model to some diagnostics. Most journals or even our supervisor will want to see the post-estimation checks carried out on your model. It will also give some level of confidence if your model passes the following tests: normality, stability, heteroscedasticity, serial correlation, model specification and so on. Regardless of your discipline, empirical model and estimation technique, it is essential that your results are supported with some “comforting” post-estimation checks. Find out those applicable to your model and technique of estimation.

14.         Forecasting or prediction/Submission
At this point, if the econometric model does not refute the theory under consideration, it may be used for predicting (forecast) future values of the dependent variable on the basis of the known or expected future values of the regressors. However, if the work is for submission, I will advise that it is proof-read as many times as possible before doing so.

I hope this step-by-step guide gives you some level of confidence to engage in research and data analysis. Let me know if you have any additions or if I omitted some salient points.

Post your comments and questions….