Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Center for AI Safety · @centerforaisafety
Words
4,786
Runtime
32:25
Speaking pace
148wpm
Reading time
20min
148 words per minute, below the 160 25th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
in this lecture we'll look at various accident models we'll start with some of the basic ones like fmea and turn to swiss cheese and bow tie then we'll have a background in complex systems to understand a more contemporary accident model that of stamp these models are theoretical constructs that help structure our reasoning they aren't used for computing and calculating specific risk estimates like in the previous lecture instead they provide a
74 words, the words spoken in the first 30 seconds at 148 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 1 |
| Average words per sentence | 4786.0 |
| Longest sentence | 4,786 words |
| Questions asked | 0 |
| Sentences containing a number | 1 |
Most used terms
Filler phrases
22 in total: like 7 · sort of 6 · actually 5 · basically 2 · uh 1 · um 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
No Script X-ray for this video: YouTube shows a Most replayed graph only once a video has enough views.
in this lecture we'll look at various accident models we'll start with some of the basic ones like fmea and turn to swiss cheese and bow tie then we'll have a background in complex systems to understand a more contemporary accident model that of stamp these models are theoretical constructs that help structure our reasoning they aren't used for computing and calculating specific risk estimates like in the previous lecture instead they provide a paradigm and a lens for understanding real world hazards and events and thinking about ways in which systems can go awry let's turn to some of these simplistic cause and effect accident models these models oversimplify the situation but they bring to light important considerations and that's why we'll look at them an early model is the failure modes and effects analysis model it involves cataloging lots of failure modes there are severities the occurrence probabilities of those failure modes detection probabilities and root causes and so on the first step is to identify various failure modes then we need to identify the potential adverse effects from those failure modes for each effect we need to identify how severe that failure or that effect can be then we need to look back at what are some potential root causes of that failure mode for each of these root causes we need to estimate what's the probability of it actually happening then we need to identify various controls and anomaly indicators so what are ways that we can work against these root causes and what are ways that we can detect whether the failure mode occurs or whether the root cause occurs so as you can see even in some basic accident analysis detection and anomaly detection is fairly integral to performing accident analysis then with these severities occurrence probabilities and detectability we can calculate the risk priority then we triage using the calculated risk priorities i should note that this concept of root causes is something we'll talk about later in this lecture basically this root cause idea can oversimplify and leave a lot of the relevant considerations out of the picture so we'll turn to this later but note that this is a a limitation of fmea let's now turn to the swiss cheese accident model the basic idea behind this model is defense in depth or that we should use multiple layers of safety barriers to achieve higher safety consider for example the hazards posed by viruses to reduce the risks from this we could avoid large indoor gatherings or maintain social distance wear a mask or wash our hands if one of the layers of defense doesn't work or was insufficient for stopping the hazard perhaps one of the other layers could this is how layers can work together to ultimately cut down risk we're then not putting all of our eggs in one basket and just relying on avoiding indoor gatherings instead by doing multiple of them we're achieving higher safety for machine learning there's a similar idea we could have multiple layers of defense through different research areas one layer itself may not be enough for example we could recognize that the machine learning system that's not aligned with human values may be unsafe in and of itself but if if it is aligned that still may not be enough for example there could be black swan events that would cause ml systems to misgeneralize and pursue incorrect goals or malicious actors could launch adversarial attacks and compromise the software on which the machine learning system is running and humans may need to monitor for emergent behavior or the malicious use of ml systems this is how safety is more than just one layer of defense consequently pursuing multiple safety research avenues creates multiple layers of protection which mitigates hazards and makes machine learning systems ultimately safer let's now turn to the bow tie model the bow tie model raises an important distinction that of the distinction between preventative barriers and protective barriers a preventative barrier prevents initiating hazardous events so it decreases the probability of a hazardous event and a protective barrier minimizes the hazardous events consequences so decreases the impact of the event you might remember in the previous lecture there is a discussion of risk as the probability of a hazardous event multiplied by its impact so the preventative barriers goes after the probability and the protected barriers goes after the impact if we have preventative barriers that can help reduce risk and likewise for protected barriers in machine learning here are some examples of preventative barriers and protected barriers you can see that preventative barriers can be associated with the word proactive or could be referred to as control measures and protective barriers could be referred to as reactive barriers or recovery measures these two work together to mitigate adverse consequences from various hazards for example let's consider the hazard of proxy gaming if a model is trying to game a proxy or cheat well we could potentially reduce the probability of this happening in the first place through adversarial robustness if we make our objectives less gameable then proxy gaming is less likely to happen but even if it does happen then we have some protective barriers such as anomaly detection we could detect that there's some over-optimization going on or that there's some unusual behavior happening this allows us to step in and minimize the adverse consequences as another example of a hazard we might be concerned about our seeking ai we could prevent them from wanting to gain power in the first place with a power penalty but if that's insufficient if there's still enough pressures for it to want to seek power then we could use some protective barriers like monitoring tools to inspect the models and see and identify whenever they're exhibiting this behavior uh before it's too late so we have some protective barriers that can help minimize these consequences as well and now we get to talk about complex systems the previous accident models were a bit too simplistic but complex systems will give us vocabulary for talking about the complexities of real-world systems today we'll start upping the complexity by moving beyond linear causality linear causality describes when a single cause produces a single effect in a causal chain of events for example let's say there's a hazard or a root cause and that triggers an event which deterministically triggers some consequent event and so down the chain we ultimately get an accident the accident models in the previous slides break down accidents into such a chain of events but this is a little simplistic look at the figure at the right for instance if we're talking about this system we can see that there's a lot more going on there's an inner loop and an outer loop and there's training signal affecting the agent the agent is affecting itself through its last action its action is also influencing the environment the environment is generating a reward which influences the agent meaning there's a feedback loop between them so we can see in today's interconnected system there's often a lot more going on there are multiple causes and effects there are feedback loops there's circular causation it's also the case in today's systems there are some more indirect causes that are highly relevant like this distribution of environments is going to definitely impact the agent but that's a bit farther removed these remote and indirect or diffuse causes also can't be ignored however when we're trying to model something as a chain of events it's a lot less natural to encode this type of complexity so then it often gets omitted from this type of analysis it's also the case that these linear causality stories need an initial triggering event a root cause of some sort however there's often a lot more than just one factor that's leading to the event the root cause choice is often arbitrarily done people when there's an accident might just end up blaming the human operator because that's the most convenient thing to blame however this might just be addressing a symptom of a larger underlying problem that perhaps there were no people weren't trained appropriately or there were productivity pressures so they're overworking the people they're things like that but root cause will often simplify the picture quite a bit so today it's more fruitful to ask in our complex systems that we're interfacing with in the real world it's more fruitful to ask what factors contributed to the accident rather than what's the single thing to blame what's the single component that is ultimately responsible just as linear causality can oversimplify our analysis of real world systems so too can reductionism or analytic decomposition as we're taught in many of our courses when we're analyzing systems what we ought to do is separate the system into events or components so basically break it down analyze the parts that you've broken it down into and so analyze those parts separately and then combine the results to understand the entire system sort of divide and conquer approach but this wrongly assumes that separation does not distort the system's properties it's implicitly and wrongly assuming that each part operates independently but there may be many interdependencies between the parts it implicitly assumes that the parts act the same when ammons examined singly as when acting in the hole but they might behave differently when they're all put together the parts are not subject to feedback loops or nonlinear interactions that's another oversimplifying assumption that's implicit in reductionism and interactions between the parts when they exist can be examined pairwise so this is how if you're just trying to analyze the parts and think that we can know the system by looking at its parts and once we understand the parts we understand the entire system that picture can be somewhat misleading so to combine our analysis of reductionism and linear causality the previous accident models reduce accidents to events and just consider the events often and in a chain of and in that chain of events we're assuming that hazard is a root cause of that accident instead rather than doing that we're trying to break event rather than breaking events down into cause and effect the complex systems perspective is to see events as a product of a complex interactions between parts so if somebody is asking what's the cause of that first that suggests that there's one cause and they're often looking for some simple this led to that type of story instead this complex systems approach is ha is instead saying that's not quite the right question it's what are the various contributing factors that potentially interacted so as to produce this event that's how we're seeing events not as something spurred by a root cause a limitation of reductionism is that in many systems properties emerge and they can hardly be inferred by analyzing the system's parts in isolation behind this idea of emergence here's a quote tornadoes financial collapses and human emotions aren't found in water molecules dollar bills or carbon atoms so here are some examples of emergent properties chemicals give rise to ions which are qual ion channels which are qualitatively different from chemicals those give rise to neurons which are qualitatively different from ion channels that gives rise to a brain which gives rise to thoughts you're not going to be able to suitably analyze thoughts that well if you're using the vocabulary associated with ion channels it ends up it's a higher order structure likewise here's another emergent property where it's not neces there's qualitatively different behavior if you have more of them so small amounts of uranium are fairly insignificant but if you increase the density at which they're packed then a nuclear reaction can occur so this idea is summarized with the quote more is different and for neural networks they also display emergent properties which we'll discuss later in the course but here's an example for now deep neural networks as they get larger have more parameters they can automatically learn how to perform arithmetic so the smaller ones ones with fewer parameters don't really know really don't really figure out how to perform arithmetic but as you increase their capacity then they suddenly learn how to perform arithmetic when they're doing self-supervised learning over large text corpora the basic idea behind emergence is that the whole is more than the sum of the parts and this is one way in which reductionism doesn't capture all the complexity of the real world it doesn't capture emergence now that we've defined emergence we can define what a complex system is we take complex systems to be systems with two parts one is that they have many interacting components and two they exhibit emergent collective behavior other people might define complex systems by stipulating some additional structures such as the presence of feedback loops or non-linear interactions they might require that they are adaptive or self-organize or exhibit quote-unquote scalable structure however we won't make any of these additional assumptions we'll just use the more basic definition where we're assuming emergence in many parts as it happens many systems are complex systems for example human societies are complex systems and so are financial systems but it's not just social constructs that are complex systems biological constructs such as cells are complex systems or ant colonies weather systems are also complex systems so ecosystems animal societies power grids disease ecologies social insects geophysical systems the internet is a complex system so is the human brain so are deep learning models even you are a complex system deep learning models have many of the hallmarks of complex systems for example complex systems have highly distributed functionality quite often and so do deep learning models they don't have one neuron identify whether there's a cat in the image or not instead it's the collective functioning of many neurons that helps the model identify whether there's a cat so it's more highly distributed that's to say if there are many partial concepts that are encoded redundantly and they're aggregated together it's not just a single component the functionality is the result of highly distributed parts working together there are also numerous weak non-linear connections in deep learning models which is another common property of complex systems the connection parameters in deep learning models are more often than not non-zero and these numerous weak connections are non-linear because deep learning models will have activation functions such as jellos and sigmoids which induce non-linearities it's also the case that complex systems often exhibit self-organization and deep learning models do too when we're designing deep learning models we're not saying that a neuron at this particular location needs to identify a whisker at 27 degrees instead the neuron instead the deep network self-organizes itself to minimize the loss so it's not top-down design producing the complex deep learning system instead it's bottom-up self-organization adaptivity is another common property of complex systems which is also common for many deep learning models few shot models for instance adapt to the context or their prompt and online models adapt to distribution shifts in their environment so some models are adaptive a common property of complex systems is feedback loops and deep learning models often have feedback loops too through self-play they might learn by playing against each other and in that way there's some there's a feedback loop between the model and itself human in the loop there's a feedback loop between the human and the machine learning system in an auto auto induced distribution shift the model ends up affecting the environment producing a distribution shift in the environment and that ends up affecting the observations that the model sees those are some feedback loop examples deep learning models also exhibit scalable structure which is to say that they're scaling laws and these scaling laws show that these models scale simply and consistently and a sufficient property for these deep learning systems to be complex systems is that they exhibit emergent functionality we'll speak more about emergent functionality in a later lecture but um suffice it to say that there are many numerous capabilities that are not planned and they spontaneously turn on when you're training deep networks so deep learning models exhibit many common properties of complex systems it's also the case that the pipelines for deep learning model deployment development and monitoring are complex systems and higher than that the organizations that design operate and improve these pipelines are also complex systems so they're everywhere consequently if we have some idea about how to make complex systems safer then that gives us some information about how to make deep learning systems safer too since we can view deep learning systems as complex systems we can use our understanding of complex systems to inform us about how to make deep learning systems safer we'll walk through quotes from the systems bible which is a collection of principles that generally apply to complex systems here's one such principle a complex system's failure mode cannot ordinarily be predicted from its structure and the crucial variables are discovered by accident this is implications for making deep learning systems safer one is that contemplation armchair analysis or working everything out on a whiteboard or a priori reasoning is limited in its reach you're going to have to do continual experimentation to capture the system complexity and find the relevant variables so we can't just have a blueprint for what safety looks like on a whiteboard that isn't going to actually be safe there'll be many relevant properties of the system that you're only going to find out through interacting with it continually so we need people continually interacting with deep learning systems for them to be safe we can't swoop in at the end and say here's the sort of safety blueprint follow this now that probably isn't going to capture a lot of the complexity or address all the failure modes another property is that a large system produced by expanding the dimensions of a smaller system does not behave like the smaller system so a straightforward property is that models have emergent properties but that safe small systems are not necessarily safe when scaled so if we have a safe small system and we throw more resources at it the larger system won't necessarily be safe for example if we have a model that isn't exhibiting deception but then it might start to at all when it's larger and more capable because well deception wasn't a very good strategy when it was dumber and really unable to keep it up if it gets larger more competent maybe it could actually pull off deception so that's one way in which a smaller system may be safe but a larger system might have some emergent unsafe properties another property is that a complex system that works is invariably found to have evolved from a simple system that works you're not going to come up with a complex system from scratch a very big messy one and that's suddenly going to be safe that's not how it works instead you're going to need to have a smaller system that works as a safe one and you're going to have to scale it up now that won't necessarily be safe but at least the safe large systems will have evolved from the safe small system so a subset of scaled up small systems will be safe but that certainly isn't to say that scaling up a small system will necessarily be safe when should we use these ideas from systems thinking and once we use this sort of analytic decomposition or reductionism or a priori reasoning well systems thinking can be complementary to decomposition so you could potentially use both but systems thinking can be more fruitful if the system are set under consideration has a collective function rather than being component based so if the components work together to achieve some larger function then it might make more sense to use systems thinking as opposed to some decompositional approach if the system has lots of non-linearities or stochastic properties then systems thinking may be more appropriate meanwhile if the system has linear causality or is highly deterministic perhaps a priori reasoning or whiteboard analysis is more appropriate if the system is dynamic rather than static then systems thinking might be more useful if there's a lot of connectivity between parts as opposed to it being isolated or inert then system thinking could be useful systems thinking is developed for systems that are too complex for this sort of reductive analysis because a separation into subsystems can distort the results and systems thinking is developed for systems that are too organized for statistics because too much of the underlying structure distorts the statistics it's also designed for systems that have important emergent properties so as a schematic if there's a high degree of randomness maybe one would use statistics if there's a low degree of randomness in a low degree of coupling maybe one would use analytic decomposition but when there's a higher degree of coupling it's often more appropriate to use systems thinking let's learn more about systems thinking for getting a better understanding of safety here's some factors and feedback loops in the columbia shuttle law so here's a real world system showing many different factors describing what things work for or work against system safety so let's look at the left and start with pressure and sort of go through some of these loops pressure feeds into performance pressure an increase in performance pressure ends up increasing the launch rate an increase in the launch rate increases the success rate and an increase in the success rate can increase the launch success and if we have more launch successes this feeds into expectations which ups the performance pressure an increase in performance pressure also works against the priority of safety programs which feeds into budget cuts directed toward safety which decreases system safety efforts now system safety efforts can increase the rate of safety and the rate of safety can increase the perception of safety which can feed into an increase in complacency and an increase in complacency can work against system safety efforts so an increase in system safety efforts can diffusely end up working against itself as well we can see that system safety is a pretty complicated thing and there are many interconnected interdependent parts that determine the system's overall safety we saw that there are many factors that can affect a system's overall safety we'll provide another model a hierarchical model which we'll call a socio-technical control system at the top is the government the government can affect regulators regulators can set regulations that companies need to abide by that's managed by management and they push that down to the staff the staff ends up ultimately influencing and guiding the hazardous processes this isn't to suggest that control is just top down they're also bottom-up forces for example companies can end up affecting regulators by lobbying them so there's a feedback loop between the two there are also additional factors such as public opinion which can end up affecting the government so we can see that in indirect way public opinion could potentially end up affecting the hazardous processes downstream this picture has some complications or some limitations for instance it's focusing mostly on operations it's assuming a chain of events at each level and it's assuming a root cause of some potential accidents a more modern picture is the following which is admittedly more complicated but what we can see is that the operating process is influenced by various socio-technical factors such as congress and legislatures technically-minded people will tend to focus just on the thing in the bold box the operating process but there's a lot more to making a system safe than just what's inside that box because the operating process is influenced by many other factors so if we're talking about safety we need to think more than just what's involved in the operating process when analyzing system accidents or catastrophes people often think in terms of events and likewise when thinking about longer term issues of ai people might think about events such as an ai progre aggressively pursuing the wrong objective or an ai suddenly changing its behavior when it gets the upper hand these events are relevant but there's more than that this is because not all the important information can be located or easily identified in specific events there are often systemic factors to consider for example here are some highly relevant systemic sociotechnical factors rules and regulations social pressures productivity pressures the incentive structures within an organization competition pressures safety budget and compute allocation the safety team size matters as does its quality alarm fatigue or are operators continually hearing an alarm that has a high false positive rate then they'll start to ignore it that's a relevant factor the reduction in inspection and preventative maintenance matters a lack of defense in depth matters a lack of redundancy can be a systemic factor a lack of fail-safes the safety mechanism cost is a highly relevant systemic factor and safety culture is an important systemic factor safety culture is according to mit professor nancy levison the most important factor to fix if we want to prevent future accidents so it's important not just to think in terms of events but also some of these broader systemic factors that are in the background it's important to keep them at the forefront of your thinking about safety using the various concepts presented in this lecture such as emergence and systemic factors and nonlinear causality stamp provides a systems way of thinking about safety it views safety as an emergent property of a system safety is not found an individual component it's a property of an entire system stamp views accident causes not as simple events at the start of a linear causal chain but rather the results of forces emerging from feedback loops consequently errors are viewed as symptoms of problems not necessarily best viewed as causes stamps says that system safety requires constantly monitoring the system from drifting into an unsafe state stamp also turns our attention to design choices and risk analysis that incorporates diffuse and indirect factors rather than relying on event analysis cause and effect stories and root causes stamp emphasizes improving contributing factors such as safety budget competition pressures and safety culture the stamp perspective can be used to prevent accidents in the first place there are techniques such as stpa which builds on stamp and it tries to identify issues at the design stage so stamp isn't just giving you a different story about it it can be used in practice for the purposes of this course we're just going to focus on making sure that you understand the system's way of thinking about safety though to summarize here's a juxtaposition between assumptions from the old paradigm and the system's view of safety the old paradigm says that accidents are caused by chains of directly related events and we can understand accidents by looking at chains events leading to the accident meanwhile the system's view says that accidents are actually complex processes involving the entire sociotechnical system traditional chain of events model cannot describe this process adequately the old paradigm says that safety is increased by increasing system or component reliability meanwhile the system's view says high reliability is not sufficient for safety because safety is an emergent property the old paradigm says most accidents are caused by operator error but the system's view says operator error is actually a product of the environment it's more of a symptom of a problem the old paradigm says assigning blame is less necessary to learn from and prevent accidents whereas the systems view says let's holistically understand how the system behavior contributed to the accident the old paradigm says major accidents occur from simultaneous occurrences of random events meanwhile the systems view says systems tend to migrate toward states of higher risk consequently there are different ways of interpreting systems and thinking about how safe they are there's the linear cause and effect sort of way of looking at system safety and there's a system view hopefully this lecture has helped you understand both of these perspectives
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.