Search the KHIT Blog

Showing posts sorted by date for query "you are smarter than your data". Sort by relevance Show all posts
Showing posts sorted by date for query "you are smarter than your data". Sort by relevance Show all posts

Friday, May 3, 2024

"It doesn't matter what’s true." Schemas predominantly matter in political discourse.

What?
  
   
OK, gotta admit, I'm now a major-league Brian Klaas Fan Boy in the wake of reading his book.
 

Zuck was right. "Young people are just smarter."

Well, some of them.
 
So, Brian has a Substack (of course). I subscribed (and paid, notwithstanding my prevailing dubiety; it's a wimpy authoring platform, generally). 

I ran across this today.
Schemas and the Political Brain
Understanding how cognitive shortcuts work when processing new information is crucial to understanding modern politics—and it's a facet of cognition that Republicans manipulate extremely effectively.

Republicans have battered Democrats on messaging in recent years because they intuitively understand schemas in a way that Democrats often don’t.

What is a schema?

The word schema comes from the Greek schēmat or schēma, which means “form.” But the concept in its modern usage refers to patterns of thought that provide intellectual shortcuts for processing the information we encounter in our lives. Think of it a bit like your brain’s organizational system, which structures our knowledge, old and new. To organize vast quantities of data, we need to sort everything into categories and patterns, with simplified assumptions.

The whole world we experience is, in a sense, a giant set of data. When you go for a walk, the amount of data your brain encounters is overwhelming—the shade of color on every leaf, the patterns of cracks on the sidewalk, the faces of every person you encounter, what they’re wearing, and whether they smiled at you as you went past.

It would be impossible for your brain to process and retain all that information, particularly because most of it provides little value to you. You don’t need to recall in vivid detail whether a random woman you once walked past was wearing a hat or not. As a result, the brain does a bit of intellectual triage, where most of the information we process about the world is culled and discarded. It’s a highly efficient way of dealing with and navigating an immeasurably complex world. Our brains have evolved to process information this way because it helps us survive, able to retain what matters and forget what doesn’t...
Yeah, heuristics & stuff. Type I & Type II thinking, etc. 
 
Brian sums up...
How to use schemas

The lesson, then, is not that fact-based arguments are meaningless in politics, but rather that facts are most effective when they’re nestled within a ready-made intellectual framework for how to make sense of the world. Effective political movements use facts to reinforce schemas, but they understand that the schemas are what matter most. It’s a depressing truth, but getting the right taglines, slogans, and vivid ways of presenting political opponents is often far more important than being right.

So, if Democrats want to loosen Trump’s grip on the modern GOP, then the messaging needs to match the audience, while recognizing the schemas that Trump’s voters are using to process reality. Democrats can shout from the rooftops about how Trump is a racist authoritarian tax dodger who poses an existential threat to American democracy (facts that are certainly worth repeating!). But no matter how loudly those arguments are repeated, voters in the GOP base use schemas that simply don’t have a place for those facts. They’ll ignore them, dismiss them, contort themselves with brain gymnastics until they end up in a position where they remain comfortable within their existing cognitive framework.

What’s more likely to be effective to sway the MAGA base is to brand Trump as a loser, partly because that does fit with Republican schemas, and partly because it’s a direct attack on the schema that Trump has tried to cultivate for himself for his entire life — that he’s a winner who lives in a golden penthouse, a strategic genius who always ends up on top.

If you really want to destroy someone in politics, don’t attack them with a barrage of facts and decimal points, change the fundamental way that their own supporters perceive them. The path to political victory runs, to an astonishing degree, through psychology and neuroscience. That way true power lies.
OK. lots to remark upon. Goes to a lot of prior topics here. to wit,


Having to go back and re-study this beaut.

  
Kathryn rocks.
 
This delightful person (a Twitter/X fellow-traveler) started me down this latest rabbit hole.
  
Click
to wit,
Click
2 Core Rules:
The fallibilist rule: No one gets the final say. You may claim that a statement is established as knowledge only if it can be debunked, in principle, and only insofar as it withstands attempts to debunk it.

The empirical rule: No one has personal authority. You may claim that a statement has been established as knowledge only insofar as the method to check it gives the same result regardless of the identity of the checker, and regardless of the source of the statement.

10 Commitments:
Fallibilism
The ethos that any of us might be wrong thus we strive to keep our ideas open to criticism.

Objectivity
Understanding that truth is not based on feeling, identity, or a person's/group’s lived experience.

Exclusivity
A commitment to the idea of objective reality because what lies beyond that is chaos.

Disconfirmation
A commitment that all ideas and viewpoints are subject to challenge by the community.

Accountability
A commitment to offering our ideas for critique in good faith and acknowledge when we’re wrong.

Pluralism
Valuing viewpoint diversity and a commitment to encourage, and seek it out.

Civility
A commitment to discourage and avoid personal attacks and, instead, depersonalize rhetoric.

Professionalism
A commitment to valuing the credentials and reputations of this reality-based community.

Institutionalism
A commitment to valuing and relying upon established norms and codes of conduct.

No Bullshitting
A commitment to sincerely regard the truth instead of obscuring it.
Well, "scientific method," anyone?  Likin' the "civility" reference in particular. "Pluralism?" I'm in, notwithstanding its political contentiousness of late (e.g., DEI).

These topics also have me recalling the excellent work of Zoe Chance.

More shortly...
_________
  

Wednesday, March 29, 2023

On "Data"

This data is ... These Data ARE


The Amazon blurb:
A sweeping history of data and its their technical, political, and ethical impact impacts on our world.

From facial recognition—capable of checking people into flights or identifying undocumented residents—to automated decision systems that inform who gets loans and who receives bail, each of us moves through a world determined by data-empowered algorithms. But these technologies didn’t just appear: they are part of a history that goes back centuries, from the census enshrined in the US Constitution to the birth of eugenics in Victorian Britain to the development of Google search.

Expanding on the popular course they created at Columbia University, Chris Wiggins and Matthew L. Jones illuminate the ways in which data has have long been used as
a tool tools and a weapon weapons in arguing for what is true, as well as a means of rearranging or defending power. They explore how data was were created and curated, as well as how new mathematical and computational techniques developed to contend with that those data serve to shape people, ideas, society, military operations, and economies. Although technology and mathematics are at its heart, the story of data ultimately concerns an unstable game among states, corporations, and people. How were new technical and scientific capabilities developed; who supported, advanced, or funded these capabilities or transitions; and how did they change who could do what, from what, and to whom?

Wiggins and Jones focus on these questions as they trace data’s historical arc, and look to the future. By understanding the trajectory of data—where
it they has have been and where it they might yet go—Wiggins and Jones argue that we can understand how to bend it them to ends that we collectively choose, with intentionality and purpose.
Yeah, I know. I've pretty much lost that "data are" fight. Just shows my age and pedantic irascibility. Perhaps this was just from an Amazon copywriter.
 
In the early 1990s, in the wake of more than 5 years as a programmer and SPC analyst in an Oak Ridge radioanalytical lab prior to moving to Las Vegas, I served a tenure as Technical Editor/Writer for a digital industrial diagnostics firm in West Knoxville, TN (portable FFT analyzers and related hardware and software products). I would routinely get tech paper drafts from the engineers and promptly change all the "data is" stuff to the proper "data are." My good-ole-boy redneck boss Forrest, our crew-cut mechanical engineer VP of marketing, would then red-pen my corrections back to "data is."

"It just looks wrong."

Yeah. Too bad.

I would quietly then change his edits right back, and my final cuts went off to production.

LOL: He once took issue with some of my own copy, and forbade me from henceforth using the words "affixed" and "atop"— e.g., "With the remote sensors firmly affixed atop the turbine housings..."

"They just look faggy."—Forrest.

Groan. Whatever. And, data ARE.
 
Letter to the editor, Science Magazine, April 8, 1927

Lordy.

Will let'cha know what I find in the book.

HOW DATA HAPPENED

Came onto this book via a Jill Lepore article in The New Yorker: "The Data Delusion."
…[I]imagine that all the world’s knowledge is stored, and organized, in a single vertical Steelcase filing cabinet. Maybe it’s lima-bean green. It’s got four drawers. Each drawer has one of those little paper-card labels, snug in a metal frame, just above the drawer pull. The drawers are labelled, from top to bottom, “Mysteries,” “Facts,” “Numbers,” and “Data.” Mysteries are things only God knows, like what happens when you’re dead. That’s why they’re in the top drawer, closest to Heaven. A long time ago, this drawer used to be crammed full of folders with names like “Why Stars Exist” and “When Life Begins,” but a few centuries ago, during the scientific revolution, a lot of those folders were moved into the next drawer down, “Facts,” which contains files about things humans can prove by way of observation, detection, and experiment. “Numbers,” second from the bottom, holds censuses, polls, tallies, national averages—the measurement of anything that can be counted, ever since the rise of statistics, around the end of the eighteenth century. Near the floor, the drawer marked “Data” holds knowledge that humans can’t know directly but must be extracted by a computer, or even by an artificial intelligence. It used to be empty, but it started filling up about a century ago, and now it’s so jammed full it’s hard to open.

From the outside, these four drawers look alike, but, inside, they follow different logics. The point of collecting mysteries is salvation; you learn about them by way of revelation; they’re associated with mystification and theocracy; and the discipline people use to study them is theology. The point of collecting facts is to find the truth; you learn about them by way of discernment; they’re associated with secularization and liberalism; and the disciplines you use to study them are law, the humanities, and the natural sciences. The point of collecting numbers in the form of statistics—etymologically, numbers gathered by the state—is the power of public governance; you learn about them by measurement; historically, they’re associated with the rise of the administrative state; and the disciplines you use to study them are the social sciences. The point of feeding data into computers is prediction, which is accomplished by way of pattern detection. The age of data is associated with late capitalism, authoritarianism, techno-utopianism, and a discipline known as data science, which has lately been the top of the top hat, the spit shine on the buckled shoe, the whir of the whizziest Tesla…
Read all of it.
 
BTW: "You are smarter than your data." Selected prior KHIT blog riffs.

Stay tuned...
__________
 

Wednesday, May 27, 2020

Annie Duke ROCKS!


One mitigative personal upside of our continuing "all-covid19-all-the-time" period for this Parkinson's-addled non-essential non-worker and life-long unlearner has been the recent volume of compelling books I've consumed while getting "three weeks to the gallon of gas" (and Netflix binge-watching) here in the Homeland "shire."    

No read more fun and illuminating than Annie Duke's delightful "Thinking in Bets."

I'd gotten one of my routine Amazon email book pitches. Intrigued, I clicked on the book cover link. It was offered up as a Kindle edition special, which, with my always-accruing credits (I continue to buy a ton of books), would only set me back 59 cents.

"What have you got to lose?" Nonetheless, after reading the Amazon blurb, I went, as is my custom, to first reading the negative one-star reviews, which can often be show-stoppers (afterward, I would muse "did we read the same book?").

Never before having heard of Annie Duke, I recall also having had the fleeting, snarky thought: "Oh, will the yummie Jessica Chastain play her too in The Movie" (successor to Molly's Game).

LOL.

ANNIE IN THE NEW YORKER
Annie Duke Will Beat You at Your Own Game
    
Late last year, I wrote to Annie Duke, a former professional poker player, about the possibility of profiling her. Duke, who for years was the leading female money winner in the World Series of Poker, retired from the game six years ago and has since refashioned herself as a corporate speaker and strategic consultant. She struck me as someone with a potentially unique and strange set of perspectives on gender, celebrity, and money. We spent the next few weeks engaged in a polite game of psychological warfare. I became attuned, moment by moment, to infinitesimal shifts in power and grew obsessed with the notion that she might be playing our negotiations like a card game. I’m still not sure how much of it was in my head.

At first, Duke enthusiastically agreed to be profiled, and often responded to my e-mails with smiley faces and exclamation points. She invited me to accompany her to a charity event and suggested that I come along to her brother-in-law’s birthday party. When I asked her to recommend friends and colleagues who might have insight into her career, she responded eighteen minutes later with an annotated list of twenty-seven names. It included all living members of her immediate family, her ex-husband, various professional poker players, and celebrities she has taught to play the game. Duke seemed to understand instinctively that affording a journalist access can actually be a form of self-protection: her avid participation would decrease my need to ferret out potentially unflattering material elsewhere.

Since retiring, Duke, who has four children and lives near Philadelphia, has travelled across the country delivering keynote speeches to conferences held by the likes of Citibank, Pandora, and Marriott. She has co-authored multiple gaming guides, and her first general-interest book, “Thinking in Bets: Making Smarter Decisions When You Don’t Have All the Facts,” came out in February. The book’s premise is that poker players live in a world in which “risk is made explicit” and are therefore trained to assess incoming information logically and judiciously in a way that other people are not. “A hand of poker takes about two minutes,” she writes. “Over the course of that hand, I could be involved in up to twenty decisions. And each hand ends with a concrete result: I win money or I lose money. The result of each hand provides immediate feedback on how your decisions are faring.”

Duke argues that we bet all the time: on parenting, home buying, restaurant orders. Betting is merely “a decision about an uncertain future,” and our opponents are not other people but, rather, hypothetical versions of ourselves who have chosen differently than we have. Her most urgent message is that we should all be more comfortable living with self-doubt—not for ethical reasons but for intellectual ones. Embracing uncertainty, she argues, makes you a better thinker. “Real life consists of bluffing, of little tactics of deception, of asking yourself what is the other man going to think I mean to do,” she writes, quoting John von Neumann, the father of game theory…

Read all of it.

Also buy and carefully study all of "Thinking in Bets." Not kidding.
INTRODUCTION:  Why This Isn’t a Poker Book

CHAPTER 1:  Life Is Poker, Not Chess

Pete Carroll and the Monday Morning Quarterbacks
The hazards of resulting
Quick or dead: our brains weren’t built for rationality
Two-minute warning
Dr. Strangelove
Poker vs. chess
A lethal battle of wits
“I’m not sure”: using uncertainty to our advantage
Redefining wrong


CHAPTER 2:  Wanna Bet?
Thirty days in Des Moines
We’ve all been to Des Moines
All decisions are bets
Most bets are bets against ourselves
Our bets are only as good as our beliefs
Hearing is believing
“They saw a game”
The stubbornness of beliefs
Being smart makes it worse
Wanna bet?
Redefining confidence

CHAPTER 3:  Bet to Learn: Fielding the Unfolding Future
Nick the Greek, and other lessons from the Crystal Lounge
Outcomes are feedback
Luck vs. skill: fielding outcomes
Working backward is hard: the SnackWell’s Phenomenon
“If it weren’t for luck, I’d win every one”
All-or-nothing thinking rears its head again
People watching
Other people’s outcomes reflect on us
Reshaping habit
“Wanna bet?” redux
The hard way


CHAPTER 4:  The Buddy System
“Maybe you’re the problem, do you think?”
The red pill or the blue pill?
Not all groups are created equal
The group rewards focus on accuracy
“One Hundred White Castles…and a large chocolate shake”: how accountability improves decision-making
The group ideally exposes us to a diversity of viewpoints
Federal judges: drift happens
Social psychologists: confirmatory drift and Heterodox Academy
Wanna bet (on science)?

CHAPTER 5:  Dissent to Win
CUDOS to a magician
Mertonian communism: more is more
Universalism: don’t shoot the message
Disinterestedness: we all have a conflict of interest, and it’s contagious
Organized skepticism: real skeptics make arguments and friends
Communicating with the world beyond our group


CHAPTER 6:  Adventures in Mental Time Travel
Let Marty McFly run into Marty McFly
Night Jerry
Moving regret in front of our decisions
A flat tire, the ticker, and a zoom lens
“Yeah, but what have you done for me lately?”
Tilt Ulysses contracts: time traveling to precommit
Decision swear jar
Reconnaissance: mapping the future
Scenario planning in practice
Backcasting: working backward from a positive future
Premortems: working backward from a negative future
Dendrology and hindsight bias (or, Give the chainsaw a rest)


ACKNOWLEDGMENTS
NOTES
SELECTED BIBLIOGRAPHY AND RECOMMENDATIONS FOR FURTHER READING
I was gratified to see that a lot of the books she cites are ones I own and have read. Were I still teaching "Critical Thinking" her book would be a required text.
Once something occurs, we no longer think of it as probabilistic—or as ever having been probabilistic. This is how we get into the frame of mind where we say, “I should have known” or “I told you so.” This is where unproductive regret comes from.

By keeping an accurate representation of what could have happened (and not a version edited by hindsight), memorializing the scenario plans and decision trees we create through good planning process, we can be better calibrators going forward. We can also be happier by recognizing and getting comfortable with the uncertainty of the world. Instead of living at extremes, we can find contentment with doing our best under uncertain circumstances, and being committed to improving from our experience…

One of the things poker teaches is that we have to take satisfaction in assessing the probabilities of different outcomes given the decisions under consideration and in executing the bet we think is best. With the constant stream of decisions and outcomes under uncertain conditions, you get used to losing a lot. To some degree, we’re all outcome junkies, but the more we wean ourselves from that addiction, the happier we’ll be. None of us is guaranteed a favorable outcome, and we’re all going to experience plenty of unfavorable ones. We can always, however, make a good bet. And even when we make a bad bet, we usually get a second chance because we can learn from the experience and make a better bet the next time.

Life, like poker, is one long game, and there are going to be a lot of losses, even after making the best possible bets. We are going to do better, and be happier, if we start by recognizing that we’ll never be sure of the future. That changes our task from trying to be right every time, an impossible job, to navigating our way through the uncertainty by calibrating our beliefs to move toward, little by little, a more accurate and objective representation of the world. With strategic foresight and perspective, that’s manageable work. If we keep learning and calibrating, we might even get good at it.

Duke, Annie. Thinking in Bets (pp. 230-232). Penguin Publishing Group. Kindle Edition.


Very smart woman. Lots to ponder. You will do well to watch all of it.

THE CRUX
"Once something occurs, we no longer think of it as probabilistic—or as ever having been probabilistic. This is how we get into the frame of mind where we say, “I should have known” or “I told you so.”
Annie says poker players call this "resulting." An interesting chronic problem in this time of being fashionably "data driven," and the tendency to spuriously correlate the quality of individual decisions with their singular outcomes. "Hindsight bias," in brief.

So, how does this stuff cohere with the so-called "Science of Deliberation," scientific thinking directed at accurate decisionmaking?

UPDATE: ANNIE DUKE BOOK RECOMMENDATION

She touted this one on Twitter.


I'm a couple of chapters in thus far. Very good. I can see why she recommended it.
_____________

More to come...

Friday, August 16, 2019

As football season draws nigh, a short take on being "data driven"

My latest Harpers Magazine just arrived in the mail.


Talk about "data-driven." A short snip from an excellent long-read (paywalled) entitled 'The Wood Chipper."
...One test at the [NFL evaluation] combine is more interesting [relative to all the obvious physical stuff] and says more about how we judge than all the others put together. It’s called the Wonderlic, and it was created by a Northwestern University graduate student in 1936. His name was E. F. Wonderlic. It was an I.Q. test meant to measure cognitive ability—math, language, basic reasoning. It consisted of fifty questions, with each correct answer yielding a point, fifty being a perfect score.

You can find sample Wonderlic questions on the internet:

  • Six cooks can boil 12 pots of water in four minutes. How many cooks are needed to boil 48 pots of water in four minutes?
  • A girl is 18 years old and her brother is a third her age. When the girl is 36, what will be the age of her brother?
  • What is the 18th letter of the English alphabet?
You have 12 minutes to take the test, 50 questions in 720 seconds. That was the innovation: the pressure of the ticking clock, the deadline looming. Wonderlic meant it to measure poise, not just how a person performs but how he performs under fire. It was designed for employers. He figured they’d use it to separate the execs from the mop pushers, but it was the armed forces that took it up first, especially the air forces—Army, Navy—­whose recruiters saw in it a way to find pilots. The ticking clock was thought to mimic the pressure a flier feels in combat, under the canopy when the ­MiGs close in. Fifty seconds till contact. Ten seconds. Three. Only around 2 percent of test takers even finished.
Tom Landry, the iconic leader of the Dallas Cowboys, was one of the first N.F.L. coaches to use the Wonderlic. Born in Mission, Texas, in 1924, Landry joined the Army Air Corps soon after his brother was killed in action over the North Atlantic in 1944. Landry flew thirty sorties in a B-­17 bomber and survived a crash landing. After the war, he played football at the University of Texas. A defensive back, he was elusive and fast and hit with the sort of force that wide receivers remembered years later.

Landry was taken by the Giants in the seventh round in the 1946 draft. He played seven professional seasons, the last two as a player/coach. He ran the defense opposite the offensive coordinator and future Hall of Famer Vince Lombardi. Most people remember Landry as the taciturn Texan who coached the Cowboys for 29 seasons, had 2 Super Bowl championships and 270 wins—­the face of the franchise. Though he looked as stolid as Clint Eastwood’s “Man with No Name,” Landry, in sport coat and fedora, was in fact an innovator. It was Landry who perfected the 4–­3 defense, which football fans will recognize as a standard alignment of the game: it starts with four “down” linemen, so-­called because these huge men begin each play in a three-­point stance (fingers in the turf, asses in the air), backed by three linebackers—­hence, 4–3. And Landry was one of the first N.F.L. coaches to realize the need for intelligence testing.

The game had become so complicated by the mid-­1970s—­so many formations, each requiring a read by the quarterback, which called for a series of adjustments made at the line of scrimmage, just the sort of improvisation known to combat pilots—­that Landry wanted a better way to scout for smarts. Not just speed and strength, but can a player think as he’s getting punched in the face, or concussed, or when Dick Butkus is biting his ankle at the bottom of the pile? That’s why he remembered the . . . wait . . . what’s that test they had us take during the war?

The Wonderlic had been in circulation long enough to generate a sea of data, the sort in which experts can read patterns. Via the test, they could tell you which professions attracted the smartest (and dumbest) people. ­Twenty was said to be the average score. Above forty, you’re a genius. Below ten . . . well. The highest average scores went first to systems analysts (32), then to chemists (31), and electrical engineers (30). These are your elites. Below that come the middle class, the multitude. Accountant (28). Copywriter (27). Bank teller (22). Firefighter (21), welder (17), janitor (14). Landry began giving the test to his players in the late 1970s. The rest of the league followed. It’s been a combine staple from the start, hated and feared.

Based on the Wonderlic, we know which positions are, on average, staffed by the smartest people on a football field, and which by the stupidest. You’d probably think that quarterbacks are the smartest players—­they have to run the offense, read defensive formations, and then make necessary changes—­but you’d be wrong. Offensive tackles have the top score, 26. Then centers (25), then quarterbacks (24). Running backs are said to be the dumbest, scoring an average of 16 on the Wonderlic. It would be interesting to give players the test before and after their careers; all those head blows must have an effect.

Of course, there are exceptions, outliers. Mario Manningham, a Michigan receiver, after failing multiple drug tests, lying about it, then admitting he’d lied, scored a 6 on the Wonderlic. (The scores are supposed to be confidential, but the numbers leak.) Running back Frank Gore, a probable Hall of Famer taken in the third round in 2005, scored a 6 as well. Jeff George, a physically gifted thrower who could never get it together, got a 10 on the Wonderlic, which is about as low as it gets for a quarterback. Aaron Rodgers, considered one of the smartest players because he looks brainy and played at U.C. Berkeley, scored a 35. Eli Manning, who took the Giants to two Super Bowls, scored a 39. Eric Decker, a receiver who did not compete at the combine because of an in­jury, scored an entirely unnecessary 43 on the Wonderlic (receivers average 17). Ryan Fitzpatrick, who played quarterback at Harvard, got a 48. He went to the Rams in the seventh round in 2005. Despite his nickname (Fitzmagic) and the length of his career (he’s played fourteen N.F.L. seasons), he’s been mostly mediocre, a fact that some use to discount the importance of the Wonderlic—­Fitzpatrick got a 48 and still sucks—­but that others use to prove its relevance—­If he weren’t a genius, the guy wouldn’t have lasted ten games in the N.F.L.

Linebacker Mike Mamula scored an amazing 49 on the Wonderlic (linebackers average 19). He broke or nearly broke several records at the 1995 combine, which bumped him way up in the draft. He went from a probable third rounder to a first rounder; he was taken seventh overall by the Eagles, just behind Steve ­McNair and just ahead of Warren Sapp, but lasted a mere handful of seasons and was never better than okay. Mamula is held up as an example of all that is wrong with the combine. Great in the weight room, great on the test, shitty on the field. The guy could do everything but play.

Only one prospect has ever gotten a perfect Wonderlic score: Pat ­­McInally, a Harvard wide receiver and punter who went in the fifth round in 1975 to Cincinnati, where he played ten seasons, which brings up an interesting question: Is it bad to overachieve on the Wonderlic?

General managers tend to steer clear of those who do poorly on the test and also of those who do well. Given a choice between too smart and too dumb, they’d choose too dumb every time. (Frank Gore, 6.) Anything over a 40 tends to be seen as a potential problem. Will too smart on the test mean too much thinking on the field and too much questioning in the locker room? If you’re looking at a 45, you’re looking at a guy who knows he’s smarter than the coach and who just might lead an insurrection. Some people speak of a Wonderlic sweet spot: 30 to 38, a range that would net most elite pro quarterbacks, including Andrew Luck, Tony Romo, and Colin Kaeper­nick. You want just enough intelligence to get up and down the field. Anything more is unnecessary or even a liability…
A great piece. Subscribe. Or, buy it off the stand.

The article concludes:
With all we know about the condition of retired players and the long-­term effects of concussions, maybe the real winners are those who didn’t get picked at all.
 Like, say, my grandson Keenan.


Preocious kid tennis player, USTA-ranked 43rd nationally by age 12. Four year varsity football starter in high school, subsequently recruited by more than 20 postsecondary schools, and (mercifully) opted to go Div III for his college ride (St. Olaf). We were so relieved when it was all over and he emerged unhurt.
___

A DIFFERENT AREA OF "DATA-DRIVEN" ANALYTICS:
ALL THINGS IN "MODERATION"

 "SAM"--"Sentiment Analysis Moderation," that is. Artificial Intelligence-assisted censorship. From Naked Capitalism: "Advertisers blacklisting news and other stories containing 'controversial' words..."

Stay tuned.
_____________

More to come...

Tuesday, August 6, 2019

A.I. for the masses?

What could possibly go wrong?


In my latest snailmail Science Magazine:
Bringing machine learning to the masses

Yang-Hui He, a mathematical physicist at the University of London, is an expert in string theory, one of the most abstruse areas of physics. But when it comes to artificial intelligence (AI) and machine learning, he was naïve. “What is this thing everyone is talking about?” he recalls thinking. Then his go-to software program, Mathematica, added machine learning tools that were ready to use, no expertise required. He began to play around, and realized AI might help him choose the plausible geometries for the countless multidimensional models of the universe that string theory proposes.

In a 2017 paper, He showed that, with just a few extra lines of code, he could enlist the off-the-shelf AI to greatly speed up his calculations. “I don't have to get down to the nitty gritty,” He says. Now, He says he is “on a crusade” to get mathematicians and physicists to use machine learning, and gives about 20 talks a year on the power of these new user-friendly versions.

AI used to be the specialized domain of data scientists and computer programmers. But companies such as Wolfram Research, which makes Mathematica, are trying to democratize the field, so scientists without AI skills can harness the technology for recognizing patterns in big data. In some cases, they don't need to code at all. Insights are just a drag-and-drop away. Computational power is no longer much of a limiting factor in science, says Juliana Freire, a computer scientist at New York University in New York City who is developing a ready-to-use AI tool with funding from the Defense Advanced Research Projects Agency (DARPA). “To a large extent, the bottleneck to scientific discoveries now lies with people.”…

The AI tools are more than mere toys for nonprogrammers, says Tim Kraska, a computer scientist at the Massachusetts Institute of Technology in Cambridge who leads Northstar, a machine learning tool supported by the $80 million DARPA program called Data-Driven Discovery of Models. Wade Shen, who leads the DARPA program, says the tools can outperform data scientists at building models, and they're even better with a subject matter expert in the loop.

In a demo for Science, Kraska showed how easy it was to use Northstar's drag-and-drop interface for a serious problem. He loaded a freely available database of 60,000 critical care patients that includes details on their demographics, lab tests, and medications. In a couple of clicks, Kraska created several heart failure prediction models, which quickly identified risk factors for the condition. One model fingered ischemia—a poor blood supply to the heart—which doctors know is often codiagnosed with heart failure. That was “almost like cheating,” Kraska said, so he dragged ischemia off the list of inputs and the models immediately began to retrain to look for other predictive factors.

Maciej Baranski, a physicist at the Singapore-MIT Alliance for Research & Technology Centre, says the group plans to use Northstar to explore cell therapies for fighting cancer or replacing damaged cartilage. The system will help biologists combine the optical, genetic, and chemical data they've collected from cells to predict their behavior…

The trend toward off-the-shelf AI has risks. Machine learning algorithms are often called black boxes, their inner workings shrouded in mystery, and the prepackaged versions can be even more opaque. Novices who don't bother to look under the hood might not recognize problems with their data sets or models, leading to overconfidence in biased or inaccurate results.
   
But Kraska says Northstar has a safeguard against misuse: more AI. It includes a module that anticipates and counteracts typical rookie mistakes, such as assuming any pattern an algorithm finds is statistically significant. “In the end it actually tries to mimic what a data scientist would do,” he says.
"The trend toward off-the-shelf AI has risks. Machine learning algorithms are often called black boxes, their inner workings shrouded in mystery...Novices who don't bother to look under the hood might not recognize problems with their data sets or models, leading to overconfidence in biased or inaccurate results."
I'll re-post something from last year:
___

Another "Holy Shit" book. Yikes.

ALMOST two decades ago, when I wrote the preface to my book Causality (2000), I made a rather daring remark that friends advised me to tone down. “Causality has undergone a major transformation,” I wrote, “from a concept shrouded in mystery into a mathematical object with well-defined semantics and well-founded logic. Paradoxes and controversies have been resolved, slippery concepts have been explicated, and practical problems relying on causal information that long were regarded as either metaphysical or unmanageable can now be solved using elementary mathematics. Put simply, causality has been mathematized.”

Reading this passage today, I feel I was somewhat shortsighted. What I described as a “transformation” turned out to be a “revolution” that has changed the thinking in many of the sciences. Many now call it “the Causal Revolution,” and the excitement that it has generated in research circles is spilling over to education and applications. I believe the time is ripe to share it with a broader audience.

This book strives to fulfill a three-pronged mission: first, to lay before you in nonmathematical language the intellectual content of the Causal Revolution and how it is affecting our lives as well as our future; second, to share with you some of the heroic journeys, both successful and failed, that scientists have embarked on when confronted by critical cause-effect questions.

Finally, returning the Causal Revolution to its womb in artificial intelligence, I aim to describe to you how robots can be constructed that learn to communicate in our mother tongue— the language of cause and effect. This new generation of robots should explain to us why things happened, why they responded the way they did, and why nature operates one way and not another. More ambitiously, they should also teach us about ourselves: why our mind clicks the way it does and what it means to think rationally about cause and effect, credit and regret, intent and responsibility…


Pearl, Judea; Mackenzie, Dana. The Book of Why: The New Science of Cause and Effect (Kindle Locations 47-61). Basic Books. Kindle Edition.
This one is gonna be fun. Stay tuned. From the Atlantic interview article:
...as Pearl sees it, the field of AI got mired in probabilistic associations. These days, headlines tout the latest breakthroughs in machine learning and neural networks. We read about computers that can master ancient games and drive cars. Pearl is underwhelmed. As he sees it, the state of the art in artificial intelligence today is merely a souped-up version of what machines could already do a generation ago: find hidden regularities in a large set of data. “All the impressive achievements of deep learning amount to just curve fitting,” he said recently...
Yeah.
"If I could sum up the message of this book in one pithy phrase, it would be that you are smarter than your data. Data do not understand causes and effects; humans do."
In short, being unreflectively "data-driven" (that fashionable tech cliche) is a both naive and a cop-out. (Note: some of this will surely go -- at least tangentially --  to the "information ethics" topic of my prior post.)
___

See also my 2018 post "Data Science?"

ERRATA

This is a hoot:
Artificial intelligence is not intelligent enough or, more exactly, not imaginative enough or creative enough to make us resign thinking. Tests for artificial intelligence are not rigorous enough. It does not take intelligence to meet the Turing test – impersonating a human interlocutor – or win a game of chess or general knowledge. You will know that intelligence is artificial only when your sexbot says, ‘No.’

Fernández-Armesto, Felipe. Out of Our Minds. University of California Press. Kindle Edition, location 7720. 
This book, wow!
The speed and reach of the computer revolution raised the question of how much further it could go. Hopes and fears intensified of machines that might emulate human minds. Controversy grew over whether artificial intelligence was a threat or a promise. Smart robots excited boundless expectations. In 1950, Alan Turing, the master cryptographer whom artificial intelligence researchers revere, wrote, ‘I believe that at the end of the century the use of words and general educated opinion will have altered so much that one will be able to speak of machines thinking without expecting to be contradicted.’ The conditions Turing predicted have not yet been met, and may be unrealistic. Human intelligence is probably fundamentally unmechanical: there is a ghost in the human machine. But even without replacing human thought, computers can affect and infect it. Do they corrode memory, or extend its access? Do they erode knowledge when they multiply information? Do they expand networks or trap sociopaths? Do they subvert attention spans or enable multi-tasking? Do they encourage new arts or undermine old ones? Do they squeeze sympathies or broaden minds? If they do all these things, where does the balance lie? We have hardly begun to see how cyberspace can change the psyche. [Ibid, location 7441]
ON THE OTHER HAND

Amazon recommended this book to me:


Only $4.99 Kindle price. 5 star reviews. I precipitously did 1-Click.

My Bad. It's awful. Reads like it was written by A.I.
INTRODUCTION 

Machine learning is one in all the quickest growing areas of technology, with far-reaching applications. This textbook is intended to give a proper introduction of machine learning, and all the algorithmic paradigms that machine learning offers, in a principled way. The book provides an intensive hypothesis of the basic concepts underlying machine learning and also the mathematical derivations that remodel these principles into practical algorithms. After a presentation of the basics of the sector, the book covers a wide range of central topics that have never been addressed by previous textbooks. These embody a discussion of the process complexity of learning and also the ideas of convexity and stability; major algorithmic paradigms together with stochastic gradient descent, neural networks, and structured output learning; and rising theoretical ideas like the PAC-Bayes approach and compression-based bounds. Designed for a starting graduate or refined student course, the text makes the elemental and algorithms of machine learning accessible to non-expert readers and pupils of arithmetics, engineering, statistics and computer science.

Samelson, Steven. Machine Learning: The Absolute Complete Beginner’s Guide to Learn and Understand Machine Learning From Beginners, Intermediate, Advanced, To Expert Concepts (pp. 1-2). Kindle Edition.
Seriously? Need I really elaborate? Got played this time.

UPDATE: HEALTH CARE AI ACROSS THE POND

Reported at TechCrunch:
The UK’s National Health Service is launching an AI lab

The UK government has announced it’s rerouting £250M (~$300M) in public funds for the country’s National Health Service (NHS) to set up an artificial intelligence lab that will work to expand the use of AI technologies within the service.

The Lab, which will sit within a new NHS unit tasked with overseeing the digitisation of the health and care system (aka: NHSX), will act as an interface for academic and industry experts, including potentially startups, encouraging research and collaboration with NHS entities (and data) — to drive health-related AI innovation and the uptake of AI-driven healthcare within the NHS.

Last fall the then new in post health secretary, Matt Hancock, set out a tech-first vision of future healthcare provision — saying he wanted to transform NHS IT so it can accommodate “healthtech” to support “preventative, predictive and personalised care”.

In a press release announcing the AI lab, the Department of Health and Social Care suggested it would seek to tackle “some of the biggest challenges in health and care, including earlier cancer detection, new dementia treatments and more personalised care”.

Other suggested areas of focus include:

  • improving cancer screening by speeding up the results of tests, including mammograms, brain scans, eye scans and heart monitoring
  • using predictive models to better estimate future needs of beds, drugs, devices or surgeries
  • identifying which patients could be more easily treated in the community, reducing the pressure on the NHS and helping patients receive treatment closer to home
  • identifying patients most at risk of diseases such as heart disease or dementia, allowing for earlier diagnosis and cheaper, more focused, personalised prevention
  • building systems to detect people at risk of post-operative complications, infections or requiring follow-up from clinicians, improving patient safety and reducing readmission rates
  • upskilling the NHS workforce so they can use AI systems for day-to-day tasks
  • inspecting algorithms already used by the NHS to increase the standards of AI safety, making systems fairer, more robust and ensuring patient confidentiality is protected
  • automating routine admin tasks to free up clinicians so more time can be spent with patients...
Have to wonder what Seamus O'Mahony would say?
_____________

More to come...

Monday, October 1, 2018

"Data Science?"

The latest fad? Last year it was profitably fashionable to add "crypto" and/or "blockchain" to one's resume or startup company name. I've alluded to the phrase "data science" in a number of prior posts, in the context of Health InfoTech. See, e.g., "Health IT: process mining and analytics for healthcare QI.

(BTW: Blockchain update.)

This (below) is a pretty good illustrative graphic of the subtopical components:


I have direct work experience in a number of these areas, but not "machine learning" nor "large scale distributed computing" (and I have some methodological concerns about the latter, which I will get to). "BPM" is "Business Process Management." We called "process mining" "operations analytics."
The allusion to "databases," one assumes, includes the critical subject of "database architectures." The heterogeneity of widely distributed "big data" (often of materially varying quality pedigree) has to be a concern. In fairness, though, my waning programmer / database architect chops are pretty old-school RDBMS comprising in-house (e.g., local server) "structured data."
By "machine learning," I assume they include "artificial intelligence," "deep learning," and "natural language processing (NLP)."

I'm reading up.


Just getting started with these, stay tuned. Looking for clear, consistent definitions at the outset, for one thing.

From the MIT book:
1. What Is Data Science? 

Data science encompasses a set of principles, problem definitions, algorithms, and processes for extracting nonobvious and useful patterns from large data sets. Many of the elements of data science have been developed in related fields such as machine learning and data mining. In fact, the terms data science, machine learning, and data mining are often used interchangeably. The commonality across these disciplines is a focus on improving decision making through the analysis of data. However, although data science borrows from these other fields, it is broader in scope. Machine learning (ML) focuses on the design and evaluation of algorithms for extracting patterns from data. Data mining generally deals with the analysis of structured data and often implies an emphasis on commercial applications. Data science takes all of these considerations into account but also takes up other challenges, such as the capturing, cleaning, and transforming of unstructured social media and web data; the use of big-data technologies to store and process big, unstructured data sets; and questions related to data ethics and regulation...

Kelleher, John D.. Data Science (MIT Press Essential Knowledge series) . The MIT Press. Kindle Edition.
From the "AI Science" book:
What is Data Science?

Data science is multidisciplinary field that relies on scientific methods, statistics and algorithms to extract meaningful insights from data. At its core, data science is all about discovering useful patterns in data that can then be presented as information to tell a story or make informed decisions. It would be noticed that data science depends on techniques from a bunch of other fields such as computer science, mathematics, statistics and business analytics. It is common for data scientists to have skills across this range. Data science can be employed to derive insights from both small and large datasets and it is often a misconception that data science is only suited to so called big data.


Morgan, Peter. Data Science from Scratch with Python: Step-by-Step Guide (Kindle Locations 337-344). AI Sciences LLC. Kindle Edition.
OK. Their Venn diagram:


Another engrossing book that I'm way deep into at the moment, written by the AI eminence Judea Pearl.


This one is a total whack upside the head.
…We live in an era that presumes Big Data to be the solution to all our problems. Courses in “data science” are proliferating in our universities, and jobs for “data scientists” are lucrative in the companies that participate in the “data economy.” But I hope with this book to convince you that data are profoundly dumb. Data can tell you that the people who took a medicine recovered faster than those who did not take it, but they can’t tell you why. Maybe those who took the medicine did so because they could afford it and would have recovered just as fast without it.

Over and over again, in science and in business, we see situations where mere data aren’t enough. Most big-data enthusiasts, while somewhat aware of these limitations, continue the chase after data-centric intelligence, as if we were still in the Prohibition era.

As I mentioned earlier, things have changed dramatically in the past three decades. Nowadays, thanks to carefully crafted causal models, contemporary scientists can address problems that would have once been considered unsolvable or even beyond the pale of scientific inquiry. For example, only a hundred years ago, the question of whether cigarette smoking causes a health hazard would have been considered unscientific. The mere mention of the words “cause” or “effect” would create a storm of objections in any reputable statistical journal.

Even two decades ago, asking a statistician a question like “Was it the aspirin that stopped my headache?” would have been like asking if he believed in voodoo. To quote an esteemed colleague of mine, it would be “more of a cocktail conversation topic than a scientific inquiry.” But today, epidemiologists, social scientists, computer scientists, and at least some enlightened economists and statisticians pose such questions routinely and answer them with mathematical precision. To me, this change is nothing short of a revolution. I dare to call it the Causal Revolution, a scientific shakeup that embraces rather than denies our innate cognitive gift of understanding cause and effect.

Pearl, Judea. The Book of Why: The New Science of Cause and Effect (pp. 6-7). Basic Books. Kindle Edition
.
"If I could sum up the message of this book in one pithy phrase, it would be that you are smarter than your data. Data do not understand causes and effects; humans do." [pg. 21]
So much for the liturgy of "Data-Driven."
Among numerous other virtues, The Book of Why provides the best explication of Bayesian Networks I've ever read. I'm already long up to speed on applications of Bayes Theorem ("base rates matter"), but Pearl's Bayesian Networks stuff is off the hook, and foundational to his compelling argument.
UPDATE

Michael Lewis' new book is out. I read it all immediately.

…in the space of a few years, the interest in data analysis went from curiosity to fad. The fetish for data overran everything from political campaigns to the management of baseball teams. Inside LinkedIn, DJ presided over an explosion of job titles that described similar tasks: analyst, business analyst, data analyst, research sci. The people in human resources complained to him that the company had too many data-related job titles. The company was about to go public, and they wanted to clean up the organization chart. To that end DJ sat down with his counterpart at Facebook, who was dealing with the same problem. What could they call all these data people? “Data scientist,” his Facebook friend suggested. “We weren’t trying to create a new field or anything, just trying to get HR off our backs,” said DJ. He replaced the job titles for some openings with “data scientist.” To his surprise, the number of applicants for the jobs skyrocketed. “Data scientists” were what people wanted to be.

Lewis, Michael. The Fifth Risk (pp. 157-158). W. W. Norton & Company. Kindle Edition.
A compelling, albeit by turns depressing and infuriating read. Highly recommended.
___

"DATA SCIENCE," STANFORD IS ON IT

sdsi.stanford.edu
I saw a presentation about this stuff given by Stanford's Carlos Bustamante last December during the Health 2.0 Technology for Precision Health conference.

From the SDSI website:
Science of Data Science

Science is experiencing simultaneous challenges and opportunities at an unprecedented rate:
  • From new sources of data, especially in large quantity and unconventional structure, often from “non-scientific” sources, such as social media;
  • From new algorithmic techniques potentially expanding greatly the ability to reason from data but whose interpretation, validity and fairness can not be established by our current statistical and computational techniques;
  • From the crucial need for scientifically valid advice on questions of the greatest importance to the future of society, of life and of the earth itself---advice that must be effectively communicated to society.
In all of these, data science is clearly central. Recent computational, statistical and other research has been of great value. Much more needs to be done, however, and with a sense of urgency.

Validity of algorithmic inferences:

Algorithmic techniques to infer patterns and structure have had exceptional success recently in many areas of practical value. They can also be important, even revolutionary, for science in many areas. Data as divergent as social media interactions on one hand and satellite or drone images on the other may provide vital results through such algorithms.

However, the scientific validity of the results can not be assumed. Conventional concepts such as random sampling of the intended population are rarely relevant. A deeper understanding of the data sources and the computations applied will be essential.

Fairness of algorithmic decisions:
Beyond the scientific validity of inferences, the use of algorithmic results to recommend practical actions raises important questions of fairness and equitable treatment. Data science needs to search for valid notions of fairness, to ensure that the results of analysis and the data-based algorithms using them are fair to all demographic and other cohorts.

Privacy and the public interest:
Huge quantities of data exist for individuals, through social media, other internet activities and databases of medical, governmental, employment and commercial records. Computational and statistical techniques are needed that satisfy both the right to privacy and society’s need to deal with important questions. Progress has been made with new approaches such as differential privacy and distributed inference on private data. Much more needs to be done given the increasing attraction of mining such data sources, with the potential risks to individual rights.

Causality:
Some of the richest sources of extensive data for scientific study are observational (“non-randomized”) data bases made available by the explosion of technology (the internet and digital records in medicine, government and business). Naive application of inferential techniques to infer causal mechanisms will be seriously misleading on such data, potentially with disastrously mistaken conclusions. Research in new statistical and computational techniques to adjust for such data sources is needed.

The reproducibility crisis:
Repeated and often highly visible incidents have highlighted failures to reproduce “scientific” conclusions; for example, frequent editorials in prestigious journals such as Science and Nature have documented and apologized for many failures to reproduce published results.
Issues of scientific and academic culture are undoubtedly part of the problem. However, the radical changes in sources of data and algorithms applied mean that the practice of data analysis has changed enormously. Data science needs to find new inferential paradigms that allow data exploration prior to the formulation of hypotheses.
SDSI on Data Science in the health care space:
Data Science for Human Health
It is clear that data science will be a driving force in transitioning the world’s healthcare systems from reactive “sick-based” care to proactive, preventive care.

First, and most importantly, data science has the power to empower the consumer, giving them more control over their own care. People can make better, more informed decisions if their care providers are able to make better, more data-based recommendations. Imagine your care provider could access your genetic information in a proactive healthcare system, measure your genetic risk for disease—not just as an individual but also as a member of a larger population—and then help you manage that risk throughout your life course.

This is the kind of personalized, patient-focused medicine that current reactive healthcare systems cannot facilitate, because they are designed to wait until things go wrong with the human body before addressing the problem, and every individual is deemed responsible for managing his/her own health and risk. In a data-based proactive healthcare system, public education could inform people of what it means to have different levels of risk. Since we all carry some level of risk (some more than others for specific diseases), individuals could be informed of their individual and collective health risks early on, enhancing control over their own health at every stage of their lifespan.

Second, data science enables more cost-effective drug discovery, helping us do the right thing for the right person. Rather than have someone trying and failing ten different drugs at great expense to the individual and the acute-based care system (not to mention worsening quality of life for the patient), data science can help us choose the right one on the first try. Although that drug in isolation is more expensive for the system, it would have been even more expensive if we didn’t have data science because that person would have had ten different things tried and failed. Additionally, data science allows us to bring things to market more quickly, because we’re not beholden to the hypothesis-driven routine.

Third, data science technologies are capable of improving patient outcomes and conditions with variable outcomes. They can capture data inputs, weed out subtypes, and distill best practices when combating disease, such as brain or other neurological cancers.

Lastly, data science technology can also reconfigure the costs associated with delivery of care by utilizing continuous data capture, analytics, and new key insights in order to inform physicians and clinicians when things have gone wrong in the human body before patients feel unwell. That understanding could then be integrated into a new model of care, which would enable early intervention, thus preventing that individual from having to go to the hospital. Recent Stanford research has begun to explore the possibilities of monitoring cardiomyopathy patients at home and monitoring children in the ER and ICU: we believe these studies are leading us toward a future of proactive, consumer-based care.

We recognize fully that technological advancement and unprecedented growth in biomedical data have created great opportunities, but they have also introduced great challenges for protecting the privacy and security of patient and other research data. We must work with stakeholders and experts in the private sector and federal agencies, such as the NIH, to promote and practice robust and proactive information-security procedures to ensure appropriate stewardship of patient and research-participant data while at the same time enabling scientific and medical advances.
Highly recommend you read all of their topical domain info.


"Ethics and Data Science?" Yeah, I'm gonna get there too. For one thing, I gotta get around to evaluating this (below).

_____________

More to come...