If you did an informal survey asking the difference between the terms “assessment” and “evaluation”, there’s a good chance you would not get consensus. We have done it and have found that some call evaluation what others would label assessment. The only consistency we found was that most instructional designers refer to evaluation as what teachers in K12 think of as assessment. This anomaly creates the need for us to define for purposes of this module what the differences are. Please understand that we are doing so only to keep things straight, not to take sides in the issue.
Click the term to reveal our definition for each term:
Click to see a comparison chart
First of all, it is probably a safe assumption that evaluation will not be a part of most informal/free choice/museum designs
Having said that, there is a possibility that your clients will be interested in summative grading/scoring (especially if you are being asked to retrofit the informal experience inside a Pk-12 school setting). The point is that some decisions will have to be made regarding this and it is probably best to actually start at the end and work backward.. to wit:
Things an Instructional Designer needs to Consider as a part of designing learning experiences (regardless of format)
- What are the expected outcomes?
- Which of the two (assessment or evaluation) are the driving force behind the ability to decide whether those goals being met?
- Does it have to be an “either-or” question? Is there room for both?
- Which Theoretical Premise(s) is/are going to be the basis for the design (i.e., which end of the Learning Theory Continuum)?
- What is the validity/reliability of that premise to aid in the process of those decisions?
- How do you assess the activity in informal situations when there may not be one correct answer?
These are only some of the questions that need to be asked. In fact, one of your assignments will be to come up with one or two more….
In this course, the central focus has been informal experiences, specifically museums. That is not because we have any sort of anti-Pk-12 bias.. it is only that it is not often we get to spend some time looking at this meaningful and significant instructional format. As you may have heard.. perhaps as much as 80% of what a child learns over his or her life actually takes place outside of the classroom. Given that much of that does not even involve any formalized approach (even museums, World’s/County Fairs etc.) all of us over our lifetimes learn and probably associates some type of self-assessed judgment with the activity. So, looking at these principles is meaningful for everyone, including PK-12 teachers.
Having said all of that, the issue at hand is.. how DO we assess whether learning is taking place?
Let’s assume for the sake of argument, that we have decided (or your client/subject matter expert (SME) has decided) that assessment plays a major role in the design and that we are dealing with less than precise metrics to measure so-called “success”.
What can be done?
As you have already been shown, it is not easy. But you have also been shown many different ways over the semester that, when cobbled together, will help us be able to make some predictions/assumptions, even though they are less than precise/finite… Their power to do so draws from their combination (the outcome is greater than the sum of their individual parts).
| Throughout the semester, we have been setting the table for you, as an instructional designer, to be able to draw some inferences as to whether learning is taking place in the informal environments we have been describing. Not to offend anyone but you may have noticed that even though school districts and state departments of education want to believe it, all the standardized testing in the world may not yield the level of precision as their proponents/advocates say or believe. There are simply too many variables… and, as you may have gleaned from any educational stats courses you have taken, the error term/significance level (.05) is rather large, and would not work if we were dealing with a cancer medicine, or building structure in civil engineering, etc. and we aren’t even talking about research design issues!
As an example, in both the Program Evaluation (EDF 6461) or Analysis & Evaluation of Instructional Technology (EME 6607) courses you are introduced to a white paper describing how that once S.E.S. Providers’ supplemental programs (those after school interventions that companies sell to districts to ‘supplement’ the after school programs to assist them with drop out prevention and remediation) are approved, they are almost never disapproved because in a court of law it is almost impossible to “prove’ that the do not work! Those who have been in education for a while realize that what is being done is about as best as can be because there are simply too many confounds that make any analyses less than precise. Having said all the above, we have plenty of tools at our disposal to make some well-thought-out and informed decisions, including statistics and probability to help us make some fairly accurate predictions. An integral part of our ‘science’ of learning is of course, our two instructional cornerstones: ASSURE and ADDIE models, both of which were borrowed from software engineering and other applied sciences. |
Dr. Jordan Ellenberg, in His Book “How Not to Be Wrong: The Power of Mathematical Thinking”, he introduces the reader to several examples of utilizing statistical probability and sampling to allow one to make informed choices regarding various topics, including how many affinity groups became savvy enough to beat the odds in several state’s Powerball Lotteries. (Some of you may have also seen the movie “21” with Kevin Spacey (based on the book “Bringing Down the House)) where a group of MIGHT students used statistics and card counting in Las Vegas. NO, we are not endorsing card counting, but simply helping you understand that these statistical measurements (such as the KR-20 model) can be our friend here, especially if we have enough data at our disposal… We could now diverge and discuss/debate the pros and cons of big data in education, but no time here… safe to say, however, that the more data you have the more power it brings to the table, and better predictions can be made…
Back to the situation at hand…. current In case you are unfamiliar with the concepts of probability and sampling, we can offer you two videos from the Annenberg Foundation:
Here’s an overview of probability and how it works.
Proper sampling allows us to measure only a part of the population to be able to make some inferences about the whole.
To fully understand sampling, we also need to understand distributions.
The key to making all of this work is what is referred to as internal validity… the better you design your model the more accurate your inferences will be… all of this is the subject of the Program and Evaluation and Analysis courses mentioned above.
The Significance of Significance (Re-Visited)
We discussed in a previous lesson that finding significance in the educational environment may not be as “significant'” as it is cracked up to be. What we were saying is that because there are so many variables and because the statistics behind educational research allow up to a 5% error term, we are really not “proving’ anything. There are those (i.e. Ellenberg, 2014) who suggest that what we most often are looking for, especially in non-precise instances, simply to be playing the percentage game and not always trying to be right but, conversely it might be enough to say we are not wrong. In this case, probability and predictability are our friends. In other words, we want to view assessment in terms of how far removed we are from a simple coin toss.
Estimating outcome assessment is not very precise (nor does it need to be) and can be viewed as a continuum:
| At zero, virtually no learning is taking place | < — 10% – 25% – 50% – 75% – 100% –> | At 100, we have “proof” that learning goals are met |
| Fifty percent represents a virtual “coin toss” |
These models the same valuation that we assign to the concept of reliability… i.e., anything over 50% is considered approaching reliability. We almost never reach 100% but are relatively satisfied. In imprecise learning situations (those in which we do not or do not wish to administer testing), we always want to be better than a coin toss and are also relatively satisfied when we can demonstrate it.
Granted, this view is an overly simplified version and a lot more statistics are involved but that is the subject of our evaluation and analysis courses. What we are demonstrating here is a simplified estimate of internal validity.
How do we approach internal consistency/validity without testing?
One effective way to approach internal validity is to base your design on as many complementary instructional design principles/premises as possible and essentially ‘cobbling’ them together … the newly formed composite is what provides the ability/power to make some predictions… (the whole is greater than the sum of its parts)
Following is a list of design processes/premises that we have already introduced in this course that will assist in the development of an assessment plan. The more that are utilized the better. It is not a complete list but is a starting point.
This view of learning is called “official” because it has evolved over a century of thinking about how schools should operate.
The following chart demonstrates the ways in which both views differ:
| In the classic view, learning is: | In the ‘official’ view, learning is: |
| continual and effortless | occasional (based on effort and hard work) |
| inconspicuous | obvious (via testing) |
| unpremeditated | intentional |
| internally motivated | externally motivated |
| less apt to be forgotten | more easily forgotten |
| actually inhibited by testing | assured only through testing |
| a social activity | an intellectual activity |